Remote sensing image semantic change detection method based on difference feature guidance

Through the semantic change detection method of remote sensing images based on differential features, using multi-scale feature extraction and difference enhancement technology, the prediction difficulty and inconsistency problems in semantic change detection of remote sensing images are solved, and more efficient change detection and semantic segmentation effects are achieved.

CN120298887APending Publication Date: 2025-07-11GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510353239.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing remote sensing image semantic change detection methods have problems such as increasing prediction difficulty, inconsistency, noise introduction, false detection and missed detection when dealing with change detection and semantic segmentation tasks, and it is difficult to effectively improve detection accuracy and efficiency.

Method used

Using a differential feature guidance method, differential enhancement is carried out through multi-scale feature extraction, differential feature calculation, feature segmentation module and bidirectional encoder to avoid branch interaction and information redundancy. The Vmamba and CAS modules are used to improve the accuracy of segmented branches, and the BDE module enhances the edge clarity of the changing areas.

Benefits of technology

It improves the accuracy and efficiency of semantic change detection of remote sensing images, reduces false detection and missed detection, enhances the robustness and generalization capabilities of the model, and can more accurately identify the edges of changing areas and invariant areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298887A_ABST
    Figure CN120298887A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing image semantic change detection method based on difference feature guidance, and the method comprises the steps: obtaining difference features through employing a pixel-level subtraction method, employing a guidance segmentation branch to only focus on the segmentation of a change region, finally generating an accurate land coverage map, and guiding a change branch to generate a clear edge of the change region. According to the method, semantic segmentation is more accurate, and the edge of a change region is clearer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image semantic change detection, and more specifically to a remote sensing image semantic change detection method guided by differential features. Background Art

[0002] With the development of remote sensing technology, the importance of remote sensing image semantic change detection in the fields of resource management, environmental monitoring, urban planning, etc. has become increasingly prominent. Traditional remote sensing image semantic change detection methods mainly rely on single-branch structures, double-branch structures, and triple-branch structures to achieve change detection or semantic segmentation; among them,

[0003] The single-branch structure directly outputs a "from-to" semantic change map by fusing multi-scale features of the segmentation branch as the input of the change branch. However, this structure needs to consider all possible change combinations. For a dataset with N ground labels, the model needs to process N(N - 1) change combinations, resulting in an increase in prediction difficulty.

[0004] The double-branch structure designs two segmentation branches, which respectively output land cover maps of two time phases. However, the land cover maps output by the two branches may be inconsistent in unchanged areas and changed areas; and the interaction between the branches is insufficient, making it difficult to balance the change detection and semantic segmentation subtasks.

[0005] The triple-branch structure includes a change branch and two segmentation branches, and contains a fine-grained fusion module and a semantic information interaction strategy. The interaction strategy includes:

[0006] 1. Feeding the deep features of the segmentation branch to the change branch to retain spatial details and consider the correlation of the feature spaces of the two time-phase images.

[0007] 2. Feeding the deep features of the change branch to the segmentation branch to drive the segmentation branch to learn richer semantic consistency features.

[0008] 3. Through dual-temporal semantic correlation reasoning, enabling the deep features of the segmentation branch to be mutually transmitted to capture the change relationship across the time dimension.

[0009] Although the above strategies improve the performance of the model to a certain extent, the direct interaction of deep features may introduce noise and affect the decision-making of the model. In addition, the semantic consistency of the two time-phase images is ignored in the learning process of the segmentation branch, resulting in insufficient segmentation ability; at the same time, although the information between the two branches can be effectively combined, it increases the complexity of semantic information utilization, and there are still common problems such as misdetection, missed detection, and blurred edges in the changed areas.

[0010] In summary, the existing remote sensing image semantic change detection methods still have certain limitations when dealing with change detection and semantic segmentation tasks. Therefore, it is necessary to propose a new method to better guide the semantic change detection of remote sensing images and improve the accuracy and efficiency of detection. Summary of the Invention

[0011] In view of this, the present invention provides a remote sensing image semantic change detection method based on differential feature guidance, aiming to improve the performance of change detection and semantic segmentation by improving the branch structure and its interaction strategy, so as to provide a more effective solution for the field of remote sensing image semantic change detection.

[0012] To achieve the above object, the present invention adopts the following technical solutions:

[0013] A remote sensing image semantic change detection method based on differential feature guidance includes:

[0014] Performing multi-scale feature extraction on the remote sensing images of the first time phase and the second time phase respectively;

[0015] Determining multi-scale differential features according to the multi-scale features of the remote sensing image of the first time phase and the corresponding multi-scale features in the remote sensing image of the second time phase;

[0016] Performing change perception based on the multi-scale differential features and the multi-scale features of the remote sensing image of the first time phase to obtain the first land cover map;

[0017] Performing change perception based on the multi-scale differential features and the multi-scale features of the remote sensing image of the second time phase to obtain the second land cover map;

[0018] Performing difference enhancement based on the multi-scale differential features and the multi-scale features of the remote sensing image of the first time phase and the remote sensing image of the second time phase to obtain a binary change map;

[0019] Masking the land cover map with the binary change map to obtain a semantic change map.

[0020] Furthermore, change perception is performed through a feature segmentation module, and the feature segmentation module includes a change perception unit and a feature fusion unit;

[0021] The change perception unit is used to perform change perception according to the features of each scale and the corresponding scale difference features;

[0022] The feature fusion unit is used to fuse the change perception result with the feature of the previous scale of the current scale.

[0023] Furthermore, multiple feature segmentation modules are connected in series for gradually reverse perception of the features of each scale; among them,

[0024] The change perception unit in the feature segmentation module other than the first digit performs change perception based on the output result of the previous-level feature segmentation module and the scale difference feature.

[0025] Preferably, the number of feature segmentation modules is the same as the number of scales of the features, and the last feature segmentation module only contains a change perception module.

[0026] Furthermore, the change perception unit adds the received data and then inputs it into the perception network;

[0027] The perception network includes three parallel branches;

[0028] The first branch includes a max pooling layer, a first convolutional layer, and a function activation layer;

[0029] The second branch includes an average pooling layer and a second convolutional layer;

[0030] The third branch includes a third convolutional layer and a third normalization layer,

[0031] wherein, the results output by the first branch and the second branch are concatenated and convolved, and then dot-producted with the result output by the third branch. The obtained dot-product result is added to the added feature again to obtain the change perception result.

[0032] Furthermore, the feature fusion unit adds the received data and then inputs it into the fusion network, and adds the output result of the fusion network to the added result in the feature fusion unit again to obtain the feature fusion result; wherein,

[0033] The fusion network sequentially includes a fourth convolutional layer, a first normalization layer, a fifth convolutional layer, and a second normalization layer.

[0034] Furthermore, a bidirectional encoder is used for difference enhancement, and each scale of features is matched with a bidirectional encoder to perform difference enhancement for different-scale features and corresponding-scale difference features in the bi-temporal remote sensing images.

[0035] Furthermore, the difference enhancement steps include:

[0036] Adding the scale difference feature to the corresponding-scale feature in the first-phase remote sensing image to obtain the first-scale enhanced feature;

[0037] Adding the scale difference feature to the corresponding-scale feature in the second-phase remote sensing image to obtain the second-scale enhanced feature;

[0038] Concatenating the first-scale enhanced feature and the second-scale enhanced feature to obtain the scale combined feature;

[0039] The scale merging feature is successively passed through an average pooling layer, a multi-layer perceptron, and an activation function, and then a dot product operation is performed with the scale merging feature, and the scale merging feature is self-attention weighted according to the obtained result to obtain a difference enhancement feature.

[0040] Further, multiple bidirectional encoders are connected in series successively.

[0041] A feature fusion unit is matched after each bidirectional encoder except the first one, which is used to fuse the output result of the current bidirectional encoder and the output result of the previous bidirectional encoder, and then transmit it to the feature fusion unit matched by the next bidirectional encoder; and the result output by the feature fusion unit matched by the last bidirectional encoder is used as the finally obtained binary change map.

[0042] The remote sensing image semantic change detection method based on difference feature guidance provided by the present invention, compared with the prior art, adopts a framework of pixel-level subtraction instead of a fusion module, avoiding branch interaction; avoiding a carefully designed fusion module and complex information interaction, and avoiding information redundancy caused by the repeated use of semantic information between branches.

[0043] Specifically, the beneficial effects of the present application include:

[0044] 1) Introducing Vmamba as a feature extractor, Vmamba establishes long-range dependencies while maintaining linear computational complexity, improving the performance of the model.

[0045] 2) Obtaining the difference feature of the bi-temporal image through pixel-level subtraction to guide the segmentation branch to only focus on the segmentation of the changed area to finally generate an accurate land cover map and guide the change branch to generate a clear edge of the changed area.

[0046] Furthermore, the present application proposes a CAS module on the segmentation branch, using the difference feature to guide the segmentation branch to focus on the segmentation of the changed area to generate a land cover map, making the semantic segmentation more accurate and alleviating the problems of false detection and missed detection.

[0047] On the change branch, BDE is proposed to use the difference feature to enhance the difference between the changed area and the unchanged area, as well as the representation of the multi-scale bi-temporal change feature, so that the edge of the changed area is clearer. Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0049] Figure 1 Schematic diagram of the process of a remote sensing image semantic change detection method guided by differential features;

[0050] Figure 2 Schematic diagram of the data flow in the segmentation branch;

[0051] Figure 3 Schematic diagram of the structure of the change perception unit;

[0052] Figure 4 Schematic diagram of the structure of the feature fusion unit;

[0053] Figure 5 Schematic diagram of the data flow in the change branch;

[0054] Figure 6 Schematic diagram of the structure of the bidirectional encoder;

[0055] Figure 7 Visualization result display diagram of different methods on the SECOND dataset;

[0056] Figure 8 Visualization result display diagram of different methods on the Landsat - SCD dataset;

[0057] Figure 9 Visualization comparison diagram of the ablation experiment on the SECOND dataset. Specific implementation manner

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] The embodiments of the present invention disclose a remote sensing image semantic change detection method guided by differential features. The core idea is to use differential features to guide the learning of two segmentation branches and one change branch, facilitating the model to classify in the change area and distinguish between the change and non - change areas.

[0060] The detection method of this application, as Figure 1 , includes the following steps:

[0061] Obtain the remote sensing images of the first time phase and the second time phase respectively for multi - scale feature extraction;

[0062] According to the multi - scale features of the remote sensing image of the first time phase and the corresponding multi - scale features in the remote sensing image of the second time phase, determine the multi - scale differential features;

[0063] Perceive changes based on multi-scale differential features and the multi-scale features of the first-phase remote sensing image to obtain the first land cover map;

[0064] Perceive changes based on multi-scale differential features and the multi-scale features of the second-phase remote sensing image to obtain the second land cover map;

[0065] Enhance the differences based on multi-scale differential features and the multi-scale features of the first-phase and second-phase remote sensing images to obtain a binary change map;

[0066] Use the binary change map to mask the land cover map to obtain a semantic change map.

[0067] Embodiment 1

[0068] In this embodiment, the pre-trained VMamba-Small is used as the encoder to extract the multi-scale features of the dual-time images. Among them, the encoder has four stages, and the features of the first-phase remote sensing image extracted are The features of the second-phase remote sensing image are

[0069] Subsequently, according to the multi-scale features of the first phase and the corresponding multi-scale features in the second-phase remote sensing image, determine the multi-scale differential features, that is, And Perform pixel-level subtraction correspondingly to obtain the multi-scale difference feature D i (i = 1, 2, 3, 4).

[0070] Embodiment 2

[0071] Furthermore, use the multi-scale difference features to guide the two segmentation branches to only focus on the class classification of the changed areas, so as to prevent the phenomenon that the same type of ground object has different labels from bringing confusion to the model training, thereby generating a more accurate ground object cover map.

[0072] In this embodiment, the segmentation branch performs change perception through the feature segmentation module, and the feature segmentation module includes a change perception unit CAS and a feature fusion unit Fusionblock;

[0073] The change perception unit CAS is used to perform change perception according to the features of each scale and the corresponding scale difference features;

[0074] The feature fusion unit Fusionblock is used to fuse the change perception results with the features of the previous scale of the current scale. The feature fusion enhances the recognition and localization capabilities of the deep learning model for multi-scale targets by combining the details and semantic information of different scales, and at the same time improves the training efficiency and model robustness.

[0075] According to an embodiment of the present invention, a plurality of feature segmentation modules are connected in series, and are used to perform reverse perception of each scale feature step by step; wherein,

[0076] The change perception unit in the feature segmentation module other than the first one performs change perception based on the output result of the previous feature segmentation module and the scale difference feature.

[0077] Preferably, the number of feature segmentation modules is the same as the number of feature scales, and the last feature segmentation module only includes a change perception module.

[0078] A specific embodiment, referring to Figure 2 , Figure 2 is a structural design diagram of the segmentation branch; each segmentation branch is composed of multiple feature segmentation modules. In this embodiment, the number of feature segmentation modules is the same as the number of feature scales, that is, 4;

[0079] Taking the first segmentation branch as an example, since the multi-scale features are perceived step by step in reverse, the change perception unit in the first feature segmentation module directly receives The corresponding D4 is used to perceive the changes, and the results are input into the feature fusion unit to combine Get the first perception result

[0080] Furthermore, in the present application, multiple feature segmentation modules are connected in series, so starting from the second feature segmentation module, the output result of the previous feature segmentation module and the different scale difference features are used for change perception;

[0081] According to this embodiment, the second feature segmentation module is based on Change perception with D3, the results are similar to The fusion output is performed; the remaining feature segmentation modules are similar.

[0082] It should be noted that, for the last feature segmentation module, since there is no next-level feature to be fused, it can be set to include only a change perception unit without continuing to set a feature fusion unit.

[0083] In one embodiment, the structure diagram of the change sensing unit is shown in FIG. Figure 3 Continuing with the first segmentation branch as an example, when the change perception unit receives the scale feature Or the perception result of the previous feature segmentation module When first i Add them together and get Then it is input into the perception network; the addition operation makes the feature values of the unchanged regions approach 0, while making the feature values of the changed regions larger, so that the subsequent network classifies the unchanged regions into the same class "no-change", thus particularly focusing on the class segmentation of the changed regions and enhancing the segmentation ability of the network.

[0084] In this application, the perception network includes three parallel branches;

[0085] The first branch includes a max pooling layer, a first convolutional layer (1×1Conv), and a function activation layer (sigmoid);

[0086] The second branch includes an average pooling layer and a second convolutional layer (1×1Conv);

[0087] The third branch includes a third convolutional layer (3×3Conv) and a third normalization layer (BN),

[0088] Furthermore, the results output by the first branch and the second branch are concatenated and subjected to a convolutional operation (3×3Conv), and then dot-producted with the result output by the third branch. The obtained dot-product result is added to the original features again to obtain the change perception result As a preferred solution, the obtained dot-product result is preferably passed through 3×3Conv first, and then dot-multiplied again.

[0089] In the first branch, the max pooling layer is used to extract the key features on the feature map, reduce the spatial size of the features, and at the same time retain the most significant feature information, which helps to capture important texture and shape information. The first convolutional layer (1×1Conv) adjusts the number of channels without changing the spatial dimensions of the feature map, realizes the dimensionality reduction of the features, reduces the computational complexity, and can also perform cross-channel feature fusion. The function activation layer (sigmoid) compresses the output of the convolutional layer between 0 and 1, providing a probability interpretation, which is very useful for subsequent feature fusion and decision-making processes.

[0090] The average pooling layer in the second branch: extracts the overall features of the region by averaging the feature map, which helps to capture the global context information, rather than just focusing on the local most significant features. The effect of the second convolutional layer (1×1Conv) is similar to that of the first branch, which is used to reduce the number of channels and integrate the feature information.

[0091] The third convolutional layer (3×3Conv) in the third branch can extract richer spatial features, including edges, corners, and textures, etc., and capture more complex feature relationships. The third normalization layer (BN) batch normalization can accelerate the training process of the model. By normalizing the feature distribution, the input of each layer maintains stable mean and variance, which helps to prevent gradient vanishing and explosion and improve the generalization ability of the model.

[0092] The results output by the first branch and the second branch are concatenated. The purpose is to fuse different types of feature information. The key features extracted by max pooling and the global features extracted by average pooling can complement each other. The concatenated features are further integrated through a convolutional operation (3×3 Conv) to reduce the dimension of the feature map while retaining important feature information and learning more complex feature representations.

[0093] The result after the convolutional operation is dot - producted with the result output by the third branch. This enables the features fused from the first and second branches to interact with the more detailed spatial features extracted by the third branch. The dot - product emphasizes the correlation between the two, thereby highlighting important changing features.

[0094] Adding the obtained dot - product result to the original features again is an operation of residual connection or feature recombination. This operation helps to retain important information in the original features while introducing the changing information obtained through the interaction. Such a fusion can enhance the model's perception of changes, improving the accuracy and robustness of detection.

[0095] Through the fusion of features extracted by different branches, the model can consider both the details and context information of the image simultaneously, which is crucial for accurate segmentation. By setting each layer to strengthen the features related to changes and suppress irrelevant information, the accuracy of segmentation is improved. At the same time, the residual connection helps the model learn more complex mappings while enhancing the training stability, the robustness, and the generalization ability of the model.

[0096] In one embodiment, referring to Figure 4 , Figure 4 is a schematic structural diagram of the feature fusion unit.

[0097] In this application, the feature fusion unit is used to fuse the output result of the corresponding change perception unit and the scale features of the previous level. The process includes first adding and and inputting the sum into the fusion network. Then, it successively passes through the fourth convolutional layer (3×3 Conv), the first normalization layer (BN), the fifth convolutional layer (3×3 Conv), and the second normalization layer (BN) to obtain an output result. Further, the output result of the fusion network is added to the sum result in the feature fusion unit again to obtain the feature fusion result. The initial addition sums the corresponding output results with the scale features of the previous stage. The purpose is to fuse feature information at different levels, combine detailed features and high-level semantic information. The second addition further integrates the features, strengthens the learned features, improves the discriminability of the features, maintains the diversity of information, and avoids losing important features during the fusion process. The fourth convolutional layer (3×3Conv) extracts local features and enhances the representational ability of the features. The first normalization layer (BN) accelerates the convergence speed of model training. The fifth convolutional layer (3×3Conv) further extracts and refines the features to help the model learn more complex feature representations. The second normalization layer (BN) performs normalization again after the second convolution to ensure the stability of the feature distribution, helps prevent gradient vanishing or explosion, and maintains the stability of training. As a preferred implementation, first use a 1×1 convolutional layer to adjust the scale of the low-level feature map to match the number of channels of the high-level feature map Then, add and combine the adjusted low-level feature map and high-level feature map at the element level. Subsequently, use the residual layer to refine and smooth the combined feature map. After the upsampling operation, these processed feature maps are passed to the next stage of the network

[0098] Embodiment III

[0099] Due to the influence of illumination and seasonal changes on remote sensing images, the existing models have insufficient ability to identify the boundaries between changed and unchanged regions, resulting in the generation of binary CM and further causing blurred SCM. Therefore, in order to overcome the ambiguity of the edges of the binary change map, this embodiment uses a bidirectional encoder BDE to perform reverse reinforcement based on different scale features and corresponding feature differences

[0100] Specifically, each scale of feature matches a bidirectional encoder BDE, which is used to enhance the differences of different scale features and corresponding scale difference features in the double-temporal remote sensing images

[0101] In one embodiment, referring to Figure 5 , there are 4 bidirectional encoders, which is consistent with the number of extracted feature scales

[0102] In this embodiment, multiple bidirectional encoders are connected in series in sequence. The first bidirectional encoder performs difference enhancement according to and as well as D4 to obtain The second bidirectional encoder directly performs difference enhancement according to and as well as D3 to obtain And so on

[0103] Further, behind each bidirectional encoder other than the first encoder in this application, a feature fusion unit (Fusion block) is matched to fuse the output result of the current bidirectional encoder and the output result of the previous bidirectional encoder and then transmit it to the feature fusion unit matched with the next bidirectional encoder. The purpose is to integrate features of different scales, achieve comprehensive learning of image details and overall structures, capture context information of different sizes, and help accurately divide the changed and unchanged regions; and use the result output by the feature fusion unit matched with the last bidirectional encoder as the final obtained binary change map.

[0104] Furthermore, in this embodiment, referring to Figure 6 , the difference enhancement step includes:

[0105] Adding the scale difference feature D i to the corresponding scale feature in the first-phase remote sensing image to obtain the first-scale enhanced feature;

[0106] Adding the scale difference feature D i to the corresponding scale feature in the second-phase remote sensing image to obtain the second-scale enhanced feature, which is used to increase the distance between the changed feature and the unchanged feature to emphasize the edge information;

[0107] Concatenating the first-scale enhanced feature and the second-scale enhanced feature to obtain the scale merged feature

[0108]

[0109] Making the scale merged feature sequentially pass through an average pooling layer, a multi-layer perceptron, and a sigmoid activation function and then perform a dot product operation with the scale merged feature so that the model can abstract higher-level feature representations from the original pixel-level features, learn higher-level features, thus having better generalization ability, and improving the calculation efficiency by reducing the feature dimension.

[0110] Further, perform self-attention weighting on the scale merged feature according to the obtained result to consider the global context information in the image, help the model capture the dependencies between distant pixels, strengthen the important changed information in the feature map, and at the same time suppress the unimportant background noise, improve the accuracy and robustness of change detection, and finally obtain the difference enhanced feature.

[0111] In this embodiment, the weighting process is as follows:

[0112] Determine the V value and K value of the attention respectively according to the dot product operation result;

[0113] and determining the Q value of attention according to the scale-combined features;

[0114] Then multiply the K value by the Q value and further perform a multiplication operation with the V value to obtain the difference-enhanced feature.

[0115] To further optimize the above technical solution, during training, the labels of the dataset have multiple semantic labels in the change region, while in the non-change region, they are uniformly labeled as "no change"; that is, the same type of ground object in a remote sensing image has two labels at the same time. The building label in the change region is "Building", and the label in the non-change region is "No Change"; to reduce omissions and false detections.

[0116] Furthermore, the method of the present invention was used to evaluate two public datasets, SECOND and Landsat-SCD. The experimental results show the effectiveness of the method. On the SECOND dataset, the method of the present invention reaches 64.86% on Fscd and 73.72% on the mean intersection over union (mIoU). In the Landsat-SCD dataset, Fscd is 92.52% and mIoU is 92.05%, exceeding the existing methods.

[0117] See specifically Figure 7 , Figure 7 for the visualization results of different methods on the SECOND dataset,

[0118] Figure 7 which shows the prediction results of DFGNet and the comparison methods on the SECOND dataset. The first two columns show the bi-temporal image of the land cover category and the corresponding GT label. DFGNet is superior to the comparison methods in identifying both the change region and the change type (as shown by the red dashed box and purple dashed box in the figure). As Figure 7 shown, among the comparison methods, the "water" category in the purple dashed box part cannot be well identified and is regarded as an undetected or misdetected area, indicating that DFGNet can make predictions for undetected or misdetected areas; the red dashed box part shows the change region division ability of the comparison methods and DFGNet. The first three comparison methods cannot identify this change region. ChangeMamba can identify most of the change regions and is the best-performing method among the comparison methods. DFGNet can also identify the corresponding change region. Compared with ChangeMamba, the edge of the change region detected by DFGNet is smoother.

[0119] Figure 8Shows the prediction results of DFGNet and the comparison methods on SECOND and Landsat - SCD. On the Landsat - SCD dataset, the ground sampling distance is relatively high, which requires the model to better preserve spatial details, such as Figure 8 shown. The first four models can only identify a small part of the "farmland" category in the purple box. The model we proposed can predict the "farmland" category more completely, which indicates that our model can accurately capture land cover types. In the part of the red dashed box, the edges of the "water" category predicted by the first three models are very blurred and even merged into a mass. The model tends to classify an area as the same category, which means that the first three models cannot clearly extract the changed areas, while our model can accurately capture fine - grained changes, making the edges of the changed areas clear.

[0120] The model is implemented using pytorch on an NVIDIA RTX3090 GPU. On the SECOND dataset, the pre - and post - temporal image pairs and related labels are cropped to 256×256 pixels and input into the network. Then, the network trained on the test set is used to infer the data of the original size. During the training process, the Adam W optimizer is used to optimize the network. Grouped learning rates are used, which are 5e - 5 and 6e - 4 respectively, and the weight decay is 5e - 3. We set the number of training iterations to 30000 times, and the batch size is 8. On the Landsat - SCD dataset, the original - size images are used in both the model training and test inference phases. The values of the grouped learning rates are 5e - 4 and 1e - 3 respectively. The number of training iterations is set to 200000 times, and the batch size is 4. The same optimizer and weight decay as those of SECOND are used. Random rotation, left - right flipping, and top - down flipping are used as training data augmentation methods for both datasets.

[0121] The visualization results are as Figure 9 shown. The blue box represents the effect of changed area extraction, and the green box represents the effect of semantic segmentation. In Figure 7 a1 and Figure 7 b1 in, the blue - boxed area in the GT is marked as an unchanged area. Compared with the base and base + CAS models without integrating the difference enhancement module, they wrongly identify the area within the blue box as a changed area. While in the base + DEM and DFGNet models, due to the addition of the difference enhancement module, the model can accurately judge the blue - boxed area as an unchanged area. This shows that the models integrating the difference enhancement module show advantages when performing change detection tasks.

[0122] Figure 7As can be seen from a2, when comparing the base and base+DEM models without an integrated feature segmentation module, there will be cases of large-scale category segmentation errors, misjudging "groud" as "building". However, in the base+CAS and DFGNet models, after adding the feature segmentation module, the scope of misjudgment is reduced. This shows that the models with an integrated feature segmentation module have stronger category segmentation ability in changing regions.

[0123] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method section.

[0124] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A semantic change detection method for remote sensing images guided by differential features, characterized in that Multiscale feature extraction is performed on the remote sensing images of the first and second phases respectively; According to the multiscale features of the remote sensing image of the first phase and the corresponding multiscale features in the remote sensing image of the second phase, multiscale differential features are determined; Change perception is performed based on the multiscale differential features and the multiscale features of the remote sensing image of the first phase to obtain the first land cover map; Change perception is performed based on the multiscale differential features and the multiscale features of the remote sensing image of the second phase to obtain the second land cover map; Difference enhancement is performed based on the multiscale differential features and the multiscale features of the remote sensing images of the first and second phases to obtain a binary change map; The land cover map is masked by the binary change map to obtain a semantic change map.

2. The remote sensing image semantic change detection method according to claim 1, wherein, Change perception is performed through a feature segmentation module, and the feature segmentation module includes a change perception unit and a feature fusion unit; The change perception unit is used to perform change perception according to the features of each scale and the corresponding scale difference features; The feature fusion unit is used to fuse the change perception result with the feature of the previous scale level of the current scale.

3. The remote sensing image semantic change detection method according to claim 2, wherein, There are multiple feature segmentation modules connected in series, which are used for gradually reverse perception of the features of each scale; among them, The change perception unit in the feature segmentation module other than the first and last ones performs change perception based on the output result of the previous feature segmentation module and the scale difference features.

4. The remote sensing image semantic change detection method according to claim 2 or 3, characterized in that The change perception unit adds the received data and then inputs it into the perception network; The perception network includes three parallel branches; The first branch includes a max pooling layer, a first convolutional layer, and a function activation layer; The second branch includes an average pooling layer and a second convolutional layer; The third branch includes a third convolutional layer and a third normalization layer, Among them, the results output by the first branch and the second branch are concatenated and convolved, and then dot-producted with the result output by the third branch. The obtained dot-product result is added to the added feature again to obtain the change perception result.

5. The remote sensing image semantic change detection method according to claim 2, characterized in that, The feature fusion unit adds the received data and then inputs it into the fusion network, and adds the output result of the fusion network to the added result in the feature fusion unit again to obtain the feature fusion result; among them, The fusion network sequentially includes a fourth convolutional layer, a first normalization layer, a fifth convolutional layer, and a second normalization layer.

6. The remote sensing image semantic change detection method according to claim 1, characterized in that A bidirectional encoder is used for difference enhancement, and each scale of features is matched with a bidirectional encoder, which is used for difference enhancement of different scale features and corresponding scale difference features in the dual-phase remote sensing images.

7. The remote sensing image semantic change detection method according to claim 1 or 6, characterized in that, The difference enhancement steps include: Adding the scale difference features to the corresponding scale features in the remote sensing image of the first phase to obtain the first scale enhanced feature; Adding the scale difference features to the corresponding scale features in the remote sensing image of the second phase to obtain the second scale enhanced feature; Concatenating the first scale enhanced feature and the second scale enhanced feature to obtain a scale combined feature; Making the scale combined feature pass through an average pooling layer, a multi-layer perceptron, and an activation function in sequence, and then performing a dot-product operation with the scale combined feature, and performing self-attention weighting on the scale combined feature according to the obtained result to obtain the difference enhanced feature.

8. The remote sensing image semantic change detection method according to claim 6, wherein, Multiple bidirectional encoders are connected in series in sequence, A feature fusion unit is matched to each of the bidirectional encoders except the first one. After fusing the output result of the current bidirectional encoder and the output result of the previous bidirectional encoder, it is transmitted to the feature fusion unit matched by the next bidirectional encoder; and the result output by the feature fusion unit matched by the last bidirectional encoder is used as the finally obtained binary change map.

Citation Information

Cited By

  • Image fusion method and device, storage medium and program product

    CN120543395A

  • Image change detection method and device based on state space, equipment and medium

    CN120612496A