Multi-scale feature fusion remote sensing image change detection method based on difference enhancement
By using differential enhancement and multi-scale feature fusion technology in remote sensing image change detection, MFM and DEM modules are built, and the problem of subtle changes recognition in complex surface environments is solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510075441.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
Existing remote sensing image change detection technologies are difficult to effectively identify subtle changes in complex surface environments, and the lack of data sets, insufficient cross-domain generalization capabilities, and difficult to meet real-time requirements.
Using a multi-scale feature fusion method based on differential enhancement, the multi-scale feature fusion module (MFM) and differential enhancement module (DEM) are constructed, and the dual-time phase feature map is integrated to enhance texture and feature representation capabilities, suppress noise interference, and improve the robustness and accuracy of change detection.
It significantly improves the accuracy and robustness of remote sensing image change detection, can more accurately identify changing features in complex scenes, and enhances the model's capture ability and detection accuracy of changing areas.
Smart Images

Figure CN119992385A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to computer and remote sensing image technology, and in particular to a multi-scale feature fusion remote sensing image change detection method based on difference enhancement. Background Art
[0002] Remote sensing image change detection is an important research topic in remote sensing technology and is widely used in environmental monitoring, urban planning, disaster assessment, and agricultural management. The main goal of remote sensing image change detection is to automatically identify and quantify land cover changes by comparing remote sensing images acquired at different times. With the continuous advancement of remote sensing technology and imaging equipment, it has become increasingly convenient to obtain high-resolution multi-temporal remote sensing images. However, the complex surface environment, diverse change types, and subtle change characteristics make change detection face many challenges.
[0003] Difference Enhancement and Multi-Scale Feature Fusion technologies provide effective solutions to key issues in change detection. The core idea of difference enhancement is to improve the robustness of change detection by amplifying the feature differences in the changed areas while suppressing the interference of the non-changed areas. Common difference enhancement technologies include difference map construction and attention mechanism, which can effectively alleviate the interference of non-changing factors such as lighting conditions and cloud occlusion; while multi-scale feature fusion focuses on capturing the multi-level expression of change information, and organically combines large-scale global changes with local detail changes through technologies such as multi-resolution convolution and deep residual connection. The combination of the two can complement each other: the difference enhancement module extracts significant change features, and the multi-scale feature fusion module captures change information at different levels, thereby improving the performance of the model in complex scenarios.
[0004] Multi-scale feature fusion technology based on difference enhancement has shown significant application potential in the field of remote sensing change detection. Compared with traditional methods, this technology can better suppress interference noise, extract significant features of the change area, and balance global and local information, thereby achieving comprehensive improvement in accuracy and robustness. This technology is not only suitable for traditional scenarios such as disaster monitoring and land use analysis, but also can cope with complex change patterns such as urban expansion and ecological protection. However, current research still faces challenges such as lack of data sets, insufficient cross-domain generalization capabilities, and real-time requirements. Future research can focus on the optimization of efficient fusion strategies, the design of more robust cross-scenario models, and the real-time application of large-scale change detection, further promoting the development of remote sensing image change detection technology. Summary of the invention
[0005] The purpose of the present invention is to provide a remote sensing image change detection method based on multi-scale feature fusion based on difference enhancement in view of the deficiencies of the prior art. This method can improve the accuracy and robustness of remote sensing image change detection.
[0006] The technical solution for achieving the purpose of the present invention is:
[0007] A multi-scale feature fusion remote sensing image change detection method based on difference enhancement includes the following steps:
[0008] 1) Constructing datasets: Comparison and ablation experiments are performed on five different datasets, namely CLCD, WHU, LEVIR-CD, NJDS, and GBCNR, to verify the effectiveness of the technical solution.
[0009] The CLCD dataset is a public farmland dataset that contains 2,400 pairs of farmland change samples taken by Gaofen. The image size is 256×256 pixels. These dual-phase images were taken in 2017 and 2019 in Guangdong Province, China. The spatial resolution of these images ranges from 0.5m to 2m. Each group of samples contains two images and a binary label representing farmland changes. All samples are divided into training set, validation set and test set in a ratio of 6:2:2. The number of samples in the training set, validation set and test set are 1,440, 480 and 480 respectively.
[0010] WHU This is a public dataset for building CDs. It contains a pair of high-resolution dual-phase aerial images with a resolution of 0.2m. The image size is 32,507×15,354 pixels. The data covers areas that have experienced earthquakes and reconstruction, including building renovations. These images are cropped into non-overlapping tiles with a resolution of 256×256 pixels. The number of samples in the training set, validation set, and test set are 5947, 744, and 744, respectively.
[0011] The LEVIR-CD dataset is a public building CD dataset captured from Google Earth. It consists of 637 pairs of dual-phase images with an image size of 1024×1024 pixels. The spatial resolution of these image pairs is 0.5m and the time span is 5 to 14 years. The dataset includes complex changes in villas, small garages, high-rise apartments, and large warehouses. All dual-phase image pairs are annotated with binary labels. The images are cropped into non-overlapping blocks of 256×256 pixels. These block pairs are randomly divided into training, validation, and test sets, with sample numbers of 7120, 1024, and 2048, respectively.
[0012] The NJDS dataset contains dual-phase images of Nanjing in 2014 and 2018. The images are from Google Earth and include different types of low-rise, mid-rise, and high-rise buildings. The images are cropped into non-overlapping blocks of 256×256 pixels and randomly divided into 540 pairs for the training set, 152 pairs for the validation set, and 1827 pairs for the test set.
[0013] The GBCNR dataset uses DJI Mavic series drones to take images at a maximum altitude of 500 meters. The size of each image is 5280×3956 pixels. Under the guidance of environmental ecology experts, the LabelMe tool is used to accurately annotate the wetlands. The images of the selected areas are cropped and resized to 256×256 pixels and randomly divided into a training set with 1748 pairs of images, a validation set with 499 pairs of images, and a test set with 249 pairs of images.
[0014] 2) Construct a multi-scale feature fusion module MFM (Multi-scale Feature Fusion Module, MFM for short) to integrate the advantages of the dual-phase feature map, compensate for the lack of spatial distribution information, and enhance the representation ability of texture and features: First, the two input feature maps and The feature maps are concatenated along the channel dimension to obtain a feature map of size 2C×H×W. Then a 1×1 convolution operation is used to compress the number of channels to C, as shown in formula (1):
[0015]
[0016] in, Represents the channel concatenation operation, Conv 1×1 is a 1×1 convolution operation for channel compression. Next, the multi-scale features are extracted by using convolution operations with convolution kernels of 3×3, 5×5, and 7×7 for the preliminary fusion features, and they are concatenated, as shown in formula (2):
[0017]
[0018] Among them, Conv k×k represents a convolution operation with a kernel size of k×k, [·] represents a channel concatenation operation, and the obtained
[0019]
[0020] Finally, a deep convolution operation is used to transform the multi-scale features F multi-scale Extract deep semantic information and combine it with preliminary features Perform residual connection to get the final output F output , as shown in formula (3):
[0021]
[0022] Among them, Conv 3×3 (Relu(Conv 3×3 (·)) represents 3×3 stacked convolution and activation function, Represents the concatenation in the channel dimension, Conv 1×1 Map the number of channels of the concatenated result back to the number of channels C of the input;
[0023] 3) Constructing the Difference Enhancement Module (DEM): The Difference Enhancement Module is used to amplify the differences between the two-phase images and highlight the subtle changes in the images. This design enables the model to perform better when processing images with rich details, including:
[0024] First, the two input feature maps and Subtract and add pixel values and take the absolute value output as F diff and F sum , as shown in formula (4) and formula (5):
[0025]
[0026] F diff and F sum Generate preliminary fusion features by splicing along the channel, as shown in formula (6):
[0027] F concat =[F diff ,F sum ] (6),
[0028] in, represents the channel concatenation operation, and then 3×3 convolution, batch normalization BN and ReLU activation function are used to transform F concat After processing, the number of channels is compressed from 64 to 32, as shown in formula (7):
[0029] F conv1 =ReLU(BN(Conv 3×3 (F concat ))) (7),
[0030] Among them, Conv 3×3 represents a convolution operation with a convolution kernel of 3×3, and BN represents a batch normalization operation. In order to retain more original feature information, the initial fusion feature map Conv 3×3 (F concat ) and F conv1The feature maps are fused through residual connections to obtain F res , then F res The output diagram of the MFM module is (5) Add and subtract pixels to further capture deep-level changes and similarity information, as shown in formula (8) and formula (9):
[0031] F diff2 =F res -F (5) (8)
[0032] F sum2 =F res +F (5) (9)
[0033] F diff2 and F sum2 The preliminary fusion features are generated by splicing along the channel, as shown in formula (10):
[0034] F final_concat =[F diff2 ,F sum2 ] (10),
[0035] in, represents the channel concatenation operation. Finally, the deep fusion feature is again processed through 3×3 convolution, batch normalization and ReLU activation to generate the final output feature F 1 , as shown in formula (11):
[0036] F 1 =ReLU(BN(Conv 3×3 (F final_concat ))) (11),
[0037] Among them, the output feature F 1 It is 32×128×128.
[0038] In order to solve the false detection problem caused by the quality problem of single-phase images and enhance the model's perception of the changed area, the MFM module of this technical solution ensures that the single-phase feature map can be further enriched with multi-scale information after passing through the Transformer module. The feature map fused by the MFM module contains higher-dimensional information, so that the model can use a richer feature space in change detection, which helps to identify changes more accurately.
[0039] In order to solve the interference problem caused by the large amount of information and complex background usually contained in remote sensing images and the influence of noise on change detection results, and at the same time for some situations where the changes are very small and difficult to detect, the difference enhancement module DEM of this technical solution can effectively filter out noise, enhance the signal of the change area, and make the change features more significant, thereby improving the algorithm's capture ability and detection accuracy of the change area.
[0040] This method can improve the accuracy and robustness of change detection in remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Schematic diagram of the overall architecture of the embodiment method;
[0042] Figure 2 This is an example diagram of a public data set in the embodiment;
[0043] Figure 3 Schematic diagram of a multi-scale feature fusion module in an embodiment;
[0044] Figure 4 Schematic diagram of the multi-scale fusion process in the embodiment;
[0045] Figure 5 is a schematic diagram of a difference enhancement module in an embodiment;
[0046] Figure 6 Schematic diagram of visualization results of five remote sensing image change detection datasets in the embodiments. DETAILED DESCRIPTION
[0047] The content of the present invention is further described below in conjunction with the drawings and embodiments, but the present invention is not limited thereto.
[0048] Example:
[0049] Reference Figure 1 , a multi-scale feature fusion remote sensing image change detection method based on difference enhancement, comprising the following steps:
[0050] 1) Construct a data set: Figure 2 As shown in the figure, five different datasets, CLCD, WHU, LEVIR-CD, NJDS and GBCNR, are used for comparison and ablation experiments to verify the effectiveness of the technical solution, among which:
[0051] The CLCD dataset is a public farmland dataset that contains 2,400 pairs of farmland change samples taken by Gaofen. The image size is 256×256 pixels. These dual-phase images were taken in 2017 and 2019 in Guangdong Province, China. The spatial resolution of these images ranges from 0.5m to 2m. Each group of samples contains two images and a binary label representing farmland changes. All samples are divided into training set, validation set and test set in a ratio of 6:2:2. The number of samples in the training set, validation set and test set are 1,440, 480 and 480 respectively.
[0052] WHU This is a public dataset for building CDs. It contains a pair of high-resolution dual-phase aerial images with a resolution of 0.2m. The image size is 32,507×15,354 pixels. The data covers areas that have experienced earthquakes and reconstruction, including building renovations. These images are cropped into non-overlapping tiles with a resolution of 256×256 pixels. The number of samples in the training set, validation set, and test set are 5947, 744, and 744, respectively.
[0053] The LEVIR-CD dataset is a public building CD dataset captured from Google Earth. It consists of 637 pairs of dual-phase images with an image size of 1024×1024 pixels. The spatial resolution of these image pairs is 0.5m and the time span is 5 to 14 years. The dataset includes complex changes in villas, small garages, high-rise apartments, and large warehouses. All dual-phase image pairs are annotated with binary labels. The images are cropped into non-overlapping blocks of 256×256 pixels. These block pairs are randomly divided into training, validation, and test sets, with sample numbers of 7120, 1024, and 2048, respectively.
[0054] The NJDS dataset contains dual-phase images of Nanjing in 2014 and 2018. The images are from Google Earth and include different types of low-rise, mid-rise, and high-rise buildings. The images are cropped into non-overlapping blocks of 256×256 pixels and randomly divided into 540 pairs for the training set, 152 pairs for the validation set, and 1827 pairs for the test set.
[0055] In this example, the GBCNR dataset uses DJI Mavic series drones to take images at a maximum altitude of 500 meters in Guangxi Binhai National Wetland Park in Guangxi Zhuang Autonomous Region. The size of each image is 5280×3956 pixels. Under the guidance of environmental ecology experts, the LabelMe tool is used to accurately annotate the wetlands. In order to meet the input requirements of the deep learning model, the images of the selected area are cropped and adjusted to 256×256 pixels and randomly divided into a training set with 1748 pairs of images, a validation set with 499 pairs of images, and a test set with 249 pairs of images;
[0056] 2) Construct a multi-scale feature fusion module MFM to integrate the advantages of the dual-phase feature map, compensate for the lack of spatial distribution information, and enhance the representation ability of texture and features: Figure 3 As shown, first, the two input feature maps and The feature maps are concatenated along the channel dimension to obtain a feature map of size 2C×H×W. Then a 1×1 convolution operation is used to compress the number of channels to C, as shown in formula (1):
[0057]
[0058] in, Represents the channel concatenation operation, Conv 1×1 is a 1×1 convolution operation for channel compression. Next, the multi-scale features are extracted by using convolution operations with convolution kernels of 3×3, 5×5, and 7×7 for the preliminary fusion features, and they are concatenated, as shown in formula (2):
[0059]
[0060] Among them, Conv k×k represents a convolution operation with a kernel size of k×k, [·] represents a channel concatenation operation, and the obtained
[0061]
[0062] Finally, a deep convolution operation is used to transform the multi-scale features F multi-scale Extract deep semantic information and combine it with preliminary features Perform residual connection to get the final output F output , as shown in formula (3):
[0063]
[0064] Among them, Conv 3×3 (Relu(Conv 3×3 (·)) represents 3×3 stacked convolution and activation function, Represents the concatenation in the channel dimension, Conv 1×1 Map the number of channels of the concatenated result back to the number of channels C of the input;
[0065] 3) Constructing the difference enhancement module DEM: The difference enhancement module is used to amplify the differences between the two-phase images and highlight the subtle changes in the images. This design enables the model to perform better when processing images with rich details, such as Figure 5 As shown, including:
[0066] First, the two input feature maps and Subtract and add pixel values and take the absolute value output as F diff and F sum , as shown in formula (4) and formula (5):
[0067]
[0068] F diff and F sum Generate preliminary fusion features by splicing along the channel, as shown in formula (6):
[0069] F concat =[F diff ,F sum ] (6),
[0070] in, represents the channel concatenation operation, and then 3×3 convolution, batch normalization BN and ReLU activation function are used to transform F concat After processing, the number of channels is compressed from 64 to 32, as shown in formula (7):
[0071] F conv1 =ReLU(BN(Conv 3×3 (F concat ))) (7),
[0072] Among them, Conv 3×3 represents a convolution operation with a convolution kernel of 3×3, and BN represents a batch normalization operation. In order to retain more original feature information, the initial fusion feature map Conv 3×3 (F concat ) and F conv1 The feature maps are fused through residual connections to obtain F res , then F res The output diagram of the MFM module is (5) Add and subtract pixels to further capture deep-level changes and similarity information, as shown in formula (8) and formula (9):
[0073] F diff2 =F res -F (5) (8)
[0074] F sum2 =F res +F (5) (9)
[0075] F diff2 and F sum2 The preliminary fusion features are generated by splicing along the channel, as shown in formula (10):
[0076] F final_concat =[Fdiff2 ,F sum2 ] (10),
[0077] in, represents the channel concatenation operation. Finally, the deep fusion feature is again processed through 3×3 convolution, batch normalization and ReLU activation to generate the final output feature F 1 , as shown in formula (11):
[0078] F 1 =ReLU(BN(Conv 3×3 (F final_concat ))) (11),
[0079] Among them, the output feature F 1 is 32×128×128, such as Figure 4 shown.
[0080] Difference-enhanced multi-scale feature fusion remote sensing image change detection experiment: In order to analyze the performance of this method and compare it with other algorithms, the four most commonly used indicators in change detection tasks are used, including precision (Precision, P), recall (Recall, R), F1-score and intersection over union (IoU);
[0081] The graphics card used in the experiment is NVIDIA GeForce RTX 3090. This method and all baseline models are implemented on GPU based on PyTorch. This method uses ResNet-50 pre-trained on ImageNet as the CNN backbone network. The Transformer module is any commonly used Transformer architecture. The image size of the dataset is 512×512 or 256×256 pixels, the batch size is 8, the initial learning rate is 1e-4, and the model parameters are optimized by the Adam optimizer. The weight decay is set to 0.01, and the dual-phase images are data enhanced. The data enhancement operations include random rotation, vertical flipping, and horizontal flipping. The model was trained 100 times, and the optimization time on the CLCD, NJDS, WHU-CD, and LEVIR-CD datasets was 0.8 hours, 6.4 hours, 3.2 hours, and 4.4 hours, respectively.
[0082] As can be seen from Table 1, the experimental results of adding the multi-scale feature fusion module (MFM) and the differential enhancement module (DEM) in the four data sets are better than the baseline network. In addition, the experimental results of adding the multi-scale feature fusion module (MFM) and the differential enhancement module (DEM) to the baseline network are the best, which fully verifies the effectiveness of these two modules:
[0083] Table 1:
[0084]
[0085] As can be seen from Table 2, compared with other methods, this method has significant improvements in recall, F1 and Iou: the recall of this method is 78.173%, F1 is 74.966%, and Iou is 61.305%. Compared with other methods, the temporary recall rate is at least 1.551 percentage points higher. Since CLCD is a public farmland dataset, it can be concluded that this method has a higher accuracy in vegetation area changes:
[0086] Table 2
[0087]
[0088] As can be seen from Table 3, this method has achieved good results in terms of recall, F1 and Iou: the recall of this method is 91.36%, F1 is 89.90%, and Iou is 82.85%. Among them, the recall of this method is the highest among other baseline methods, and the recall of Iou is also ranked second:
[0089] Table 3
[0090]
[0091] As can be seen from Table 4, this method has achieved good results in terms of recall, F1 and Iou: the recall of this method is 91.64%, F1 is 89.57%, and Iou is 82.96%. Among them, both the recall and Iou have achieved suboptimal good results:
[0092] Table 4
[0093]
[0094] As can be seen from Table 5, this method achieved the highest scores in all four indicators, among which the recall rate of this method is 79.24%, the F1 score is 69.83%, the IoU is 53.65%, and the precision is 62.42%:
[0095] Table 5
[0097]
[0098] like Figure 6As shown in the figure, this method improves the accuracy of change detection in remote sensing images. The difference enhancement module enhances the pixel difference of the image by calculating the difference of the change feature map of the dual-phase remote sensing image, so as to more accurately capture the subtle change features, and improve the resolution of feature expression and the sensitivity of change detection by feature difference calculation and aggregation; the multi-scale feature fusion module fully integrates the deep semantic information of the two stages, extracts multi-level change features through the multi-scale fusion mechanism, and can more accurately capture the change features of targets in complex scenes, especially in the simultaneous modeling of large-scale changes and local detail changes. The collaborative optimization of DEM and MFM is an important innovation of this method. DEM provides significant change features, and MFM further extracts multi-scale semantic information on this basis. The combination of the two significantly improves the detection ability of complex scenes and subtle changes.
[0099] This method is suitable for the detection of surface changes in remote sensing images, and can be widely used in disaster monitoring, land use change analysis, urban expansion monitoring, and ecological and environmental protection. It has significant advantages in change detection of high-resolution images: first, the difference enhancement module can accurately capture subtle change features and effectively suppress the interference of non-change factors such as illumination change and shadow, thereby improving the model's adaptability to complex scenes; second, the multi-scale feature fusion module fully integrates the two-stage deep semantic information, and can simultaneously capture large-scale global changes and local detail changes, making the model more flexible and comprehensive in dealing with tasks of changes at different scales; finally, the collaborative optimization of DEM and MFM further enhances the modeling ability of complex surface change patterns, ensuring the stability and accuracy of detection results in diverse scenarios. This method has broad application prospects and can be applied to disaster emergency monitoring (such as floods, earthquakes, forest fires, etc.), land use and cover change analysis (LUCC), ecological and environmental protection (such as wetland change identification), urban expansion monitoring, agricultural dynamic assessment (such as crop growth monitoring) and other fields, especially in the change detection of high-resolution remote sensing images. With the continuous growth of remote sensing data and its increasing importance in fields such as land and resources management, ecological and environmental protection, and disaster management, this method shows great potential in improving the efficiency and accuracy of change detection, and will provide important support for the further development of remote sensing technology.
Claims
1. A multi-scale feature fusion remote sensing image change detection method based on difference enhancement, characterized in that: The steps include: 1) Constructing data sets: using five different data sets including CLCD, WHU, LEVIR-CD, NJDS and GBCNR to conduct comparison and ablation experiments and verify the effectiveness of the technical solution, among which: The CLCD dataset is a public farmland dataset that contains 2,400 pairs of farmland change samples taken by Gaofen. The image size is 256×256 pixels. These dual-phase images were taken in 2017 and 2019 in Guangdong Province, China. The spatial resolution of these images ranges from 0.5m to 2m. Each group of samples contains two images and a binary label representing farmland changes. All samples are divided into training set, validation set and test set in a ratio of 6:2:
2. The number of samples in the training set, validation set and test set are 1,440, 480 and 480 respectively. WHU This is a public dataset for building CDs. It contains a pair of high-resolution dual-phase aerial images with a resolution of 0.2m. The image size is 32,507×15,354 pixels. The data covers areas that have experienced earthquakes and reconstruction, including building renovations. These images are cropped into non-overlapping tiles with a resolution of 256×256 pixels. The number of samples in the training set, validation set, and test set are 5947, 744, and 744, respectively. The LEVIR-CD dataset is a public building CD dataset captured from Google Earth. It consists of 637 pairs of dual-phase images with an image size of 1024×1024 pixels. The spatial resolution of these image pairs is 0.5m and the time span is 5 to 14 years. The dataset includes complex changes in villas, small garages, high-rise apartments, and large warehouses. All dual-phase image pairs are annotated with binary labels. The images are cropped into non-overlapping blocks of 256×256 pixels. These block pairs are randomly divided into training, validation, and test sets, with sample numbers of 7120, 1024, and 2048, respectively. The NJDS dataset contains dual-phase images of Nanjing in 2014 and 2018. The images are from Google Earth and include different types of low-rise, mid-rise, and high-rise buildings. The images are cropped into non-overlapping blocks of 256×256 pixels and randomly divided into 540 pairs for the training set, 152 pairs for the validation set, and 1827 pairs for the test set. The GBCNR dataset uses DJI Mavic series drones to take images at a maximum altitude of 500 meters. The size of each image is 5280×3956 pixels. Under the guidance of environmental ecology experts, the LabelMe tool is used to accurately annotate the wetlands. The images of the selected areas are cropped and resized to 256×256 pixels and randomly divided into a training set with 1748 pairs of images, a validation set with 499 pairs of images, and a test set with 249 pairs of images. 2) Construct a multi-scale feature fusion module MFM: First, the two input feature maps and The feature maps are concatenated along the channel dimension to obtain a feature map of size 2C×H×W. Then a 1×1 convolution operation is used to compress the number of channels to C, as shown in formula (1): in, Represents the channel concatenation operation, Conv 1×1 The convolution operation is 1×1. Next, the multi-scale features are extracted by using convolution operations with kernels of 3×3, 5×5, and 7×7 for the preliminary fusion features, and they are concatenated, as shown in formula (2): Among them, Conv k×k represents a convolution operation with a kernel size of k×k, [·] represents a channel concatenation operation, and the obtained Finally, a deep convolution operation is used to transform the multi-scale features F multi-scale Extract deep semantic information and combine it with preliminary features Perform residual connection to get the final output F output , as shown in formula (3): Among them, Conv 3×3 (Relu(Conv 3×3 (·)) represents 3×3 stacked convolution and activation function, Represents the concatenation in the channel dimension, Conv 1×1 Map the number of channels of the concatenated result back to the number of channels C of the input; 3) Construct the difference enhancement module DEM: First, the two input feature maps and Subtract and add pixel values and take the absolute value to output as F diff and F sum , as shown in formula (4) and formula (5): F diff and F sum Generate preliminary fusion features by splicing along the channel, as shown in formula (6): F concat =[F diff ,F sum ] (6), in, represents the channel concatenation operation, and then 3×3 convolution, batch normalization BN and ReLU activation function are used to concat After processing, the number of channels is compressed from 64 to 32, as shown in formula (7): F conv1 =ReLU(BN(Conv 3×3 (F concat ))) (7), Among them, Conv 3×3 represents a convolution operation with a convolution kernel of 3×3, and BN represents a batch normalization operation. The initial fusion feature map Conv 3×3 (F concat ) and F conv1 The feature maps are fused through residual connections to obtain F res , then F res The output diagram of the MFM module is (5) Add and subtract pixels to capture deep changes and similarity information, as shown in formula (8) and formula (9): F diff2 =F res -F (5) (8), F sum2 =F res +F (5) (9), F diff2 and F sum2 The preliminary fusion features are generated by splicing along the channel, as shown in formula (10): F final_concat =[F diff2 ,F sum2 ] (10), in, represents the channel concatenation operation. Finally, the deep fusion feature is again processed through 3×3 convolution, batch normalization and ReLU activation to generate the final output feature F 1 , as shown in formula (11): F 1 =ReLU(BN(Conv 3×3 (F final_concat ))) (11), Among them, the output feature F 1 It is 32×128×128.