Semantic change detection method, system and equipment of remote sensing image and medium

The STGNet network addresses the challenges of inconsistent resolution and temporal association in semantic change detection by aligning and fusing features across time segments, improving the accuracy of semantic change detection in remote sensing imagery.

CN120318702AActive Publication Date: 2025-07-15SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Patent Information

Application Number
CN202510263431.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-15
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

In the semantic change detection of remote sensing images, the prior art has problems such as small targets and blurred boundaries, inconsistent data resolution and lack of time dimension correlation, resulting in insufficient detection accuracy and inconsistent semantic regions and change regions.

Method used

The improved semantic change detection network STGNet is adopted to optimize the detection results of the detection results through the dual-channel dual-branch network structure and the dual-time refinement interactive structure, guide the fusion of feature information, and use the dual-branch network with different channel attention weights and scales.

Benefits of technology

It improves the accuracy and consistency of semantic change detection of remote sensing images, can more accurately identify surface coverage and land use changes, and enhances the feature extraction ability of high-resolution remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318702A_ABST
    Figure CN120318702A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic change detection method, system, equipment and medium for a remote sensing image, and relates to the technical field of semantic change detection, and the method comprises the steps: obtaining a plurality of dual-time remote sensing images; constructing a double-path double-branch network structure, and guiding one branch to selectively learn and fuse the feature information of the other branch through the learning of a bidirectional guiding module; the feature information of the double branches in different time periods is kept consistent in channel number and spatial resolution, an interaction strategy is adopted, double-branch network structures with different scales are used for extracting channel attention weights respectively, dynamic feature fusion is carried out based on weight information, and double-time semantic features are obtained; and performing binary change detection on the dual-time semantic features to obtain a binary change detection graph, and obtaining a final result of semantic change detection. According to the method, a two-way guiding strategy is integrated, the network can be effectively guided to be more focused on key change characteristics in a two-time-phase image, and the method also has strong robustness and adaptability under images with different resolutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic change detection, and in particular to a method, system, device and medium for detecting semantic changes of remote sensing images. Background Art

[0002] Change detection refers to identifying the difference information by observing dual-time or multi-time images. Accurate change detection can timely detect changes in surface cover, land use, and natural phenomena, and plays an important role in environmental monitoring, urban planning, disaster warning and emergency response, and resource management.

[0003] Traditional change detection mainly relies on manual visual interpretation and simple image processing technology, which is time-consuming, labor-intensive and easily affected by subjective factors, and is difficult to cope with the processing needs of large-scale, high-resolution remote sensing data. In recent years, with the successful application of deep learning in the fields of target detection and semantic segmentation, remote sensing change detection technology based on convolutional neural network (CNN) has also made significant progress. In remote sensing change detection, CNN can better process such complex, high-dimensional high-resolution remote sensing images, and has powerful feature extraction capabilities, which is conducive to accurately identifying changes in surface cover, land use, etc. This method not only greatly improves the efficiency and accuracy of change detection, reduces the need for manual intervention and subjective judgment, but also can better adapt to the complex and changeable remote sensing data environment, providing strong support for scientific research and practical applications in various fields.

[0004] The dual-branch architecture has a feature extraction part with shared weights, which can efficiently extract common features in dual-time or multi-time input images. Therefore, most researchers use the Siamese network structure for binary change detection. However, for the semantic change detection SCD task, the detection accuracy obtained by using the dual-branch architecture mostly depends on the classification results of a single branch, which is not conducive to the subsequent change detection results. Therefore, based on the dual-branch architecture, scholars proposed a multi-task learning framework, which uses two shared-weight CNN branches to extract semantic information features of different time periods, and then uses another CNN branch to perform change analysis on the obtained dual-time features, thereby improving accuracy. Despite this, these methods still face the following challenges:

[0005] In the prior art, when dealing with small targets and boundaries, due to inconsistent data resolutions or limitations in representation methods, detailed information is lost, making it difficult to capture key semantic features. Moreover, existing models only independently extract the features of a single time period from dual-time images and do not explicitly model the correlations in the time dimension, lacking the ability to capture the changing information in the time dimension. Often, the detected change regions in the semantic change detection (BCD) task are inconsistent with the semantic regions in the semantic segmentation (SS) task, resulting in contradictory prediction results for the semantic regions and the change regions and inaccurate prediction of the unchanged regions. Summary of the Invention

[0006] The purpose of the present invention is to provide a semantic change detection method, system, device, and medium for remote sensing images in view of the above deficiencies in the prior art to solve the problems in the prior art.

[0007] The present invention specifically provides the following technical solutions:

[0008] A semantic change detection method for remote sensing images includes the following steps:

[0009] Obtain multiple dual-time remote sensing images containing land cover types;

[0010] Construct an improved semantic change detection network STGNet. The improved semantic change detection network STGNet includes a dual-path and dual-branch network structure and a dual-time refinement interaction structure, and input the dual-time remote sensing images into the improved semantic change detection network STGNet for training to obtain a trained semantic change detection network STGNet;

[0011] Input the dual-time remote sensing images to be detected into the trained semantic change detection network STGNet to obtain the semantic change detection results of the remote sensing images;

[0012] Among them, inputting the dual-time remote sensing images into the improved semantic change detection network STGNet for training includes:

[0013] Obtain the features of the dual-time remote sensing images through the dual-path and dual-branch network structure, and in each branch network structure of the dual-path and dual-branch network structure, guide one branch to learn and fuse the feature information of the other branch; the feature information includes the context information and spatial detail information of the dual-time remote sensing images containing land cover types;

[0014] In the dual-time refinement interaction structure, keep the feature information of the dual branches in different time periods consistent in terms of the number of channels and spatial resolution, and based on the interaction strategy, use dual-branch network structures with different scales to separately extract the channel attention weights in the feature information, and perform feature fusion based on the channel attention weights to obtain dual-time semantic features;

[0015] Perform binary change detection on the dual-temporal semantic features to obtain a binary change detection map, and perform masking processing on the land cover type through the binary change detection map to obtain the final result of semantic change detection.

[0016] Preferably, the dual-path and dual-branch network structure includes a Detail Awareness Path (DAP) and a Context Path (CP). The Context Path (CP) uses ResNet50 as the backbone network to extract semantic features, and the output resolutions are 1 / 2, 1 / 4, 1 / 8, 1 / 8, and 1 / 8 of the original resolution respectively. Bidirectional Guidance Modules (BiDs) are set between the two 1 / 8 resolutions; the Detail Awareness Path (DAP) captures the fine spatial structure and edge information in the remote sensing image, and the output resolutions are 1 / 2, 1 / 4, 1 / 4, and 1 / 4 of the original resolution respectively. Bidirectional Guidance Modules (BiDs) are set between the two 1 / 4 resolutions.

[0017] Preferably, guiding one branch to learn and fuse the feature information of the other branch includes:

[0018] Define the corresponding pixel vectors in feature maps A and B as and And perform dynamic convolution operations on the corresponding pixel vectors, where the dynamic convolution on and The processing process is as follows:

[0019]

[0020] Among them, π k represents the attention weight of the k-th linear function , g is the activation function, W and b are the weight matrix and bias vector respectively, and x represents the input feature component;

[0021] Perform batch normalization processing, element-wise multiplication and summation operations on the features, fuse the feature information from different branches, and process the feature information through the Sigmoid function to obtain the possibility σ that two pixels belong to the same object. If σ is higher than the threshold, trust from its own branch. Specifically:

[0022]

[0023] Among them, f represents the dynamic convolution operation, BN represents the batch normalization operation, Sum represents the summation operation, and Sigmod represents the activation function.

[0024] Preferably, based on the interaction strategy, using dual-branch network structures with different scales to extract the channel attention weights in the feature information includes:

[0025] Based on the T1 period, the high-level semantic features ST2 of the T2 period are activated through the sigmoid function, and the activated semantic features ST2 are multiplied by the detail features DT1 of the T1 period as weight factors, prompting the spatial detail information in the T1 period to learn the semantic features in the T2 period and obtaining cross-period spatial detail information F T1 ; The specific expression is:

[0026] F T1 = sigmod(DWConv 3×3 (Up(S T1 )))×DWConv 3×3 (D T1 );

[0027] Among them, DWConv 3x3 represents depthwise separable convolution with a convolution kernel size of 3x3, and Up represents upsampling.

[0028] Preferably, the feature fusion based on the channel attention weight to obtain the dual-time semantic features includes:

[0029] Performing linear addition feature fusion on the input spatial detail information X and semantic information Y to obtain a preliminary fusion result F, and the specific expression is as follows:

[0030] F = X + Y

[0031] Using global average pooling and pointwise convolution to extract the global channel attention and local channel attention in F respectively, normalizing the extracted attention weights to between 0 and 1 through the sigmoid function, and multiplying and adding the attention weights with the corresponding features to obtain the final fusion result, and the specific expression is:

[0032] L(F) = B(PWConv2(δ(B(PWConv1(F)))));

[0033] G(F) = B(PWConv2(δ(B(PWConv1(GAP(F))))));

[0034]

[0035] Among them, F represents the initial feature fusion result, PWConv1 and PWConv2 both represent 1x1 pointwise convolution, B represents the BatchNorm layer, δ represents the ReLU activation function, and GAP represents the global average pooling operation.

[0036] Preferably, when training the dual-time remote sensing images by inputting them into the improved semantic change detection network STGNet, it further includes:

[0037] Adopt the semantic loss function \(L\) s and the binary change function \(L\) c and the semantic change loss function \(L\) SC to optimize the semantic change detection results and obtain the optimized semantic change detection network STGNet; specifically:

[0038] The semantic loss function \(L\) s performs multi-class cross-loss entropy calculation on the semantic classification results \(S1\) and \(S2\) and the semantic categories \(TS1\) and \(TS2\) in the true value semantic change map, and is specifically expressed as:

[0039]

[0040] where \(y\) i and respectively represent the true semantic label category and the probability predicted as the \(i\)-th category, and \(N\) represents the land cover types in semantic change detection;

[0041] The binary change function \(L\) c obtains the binary cross-entropy loss between the predicted change map and the true change map, and is specifically expressed as:

[0042]

[0043] where \(y\) c represents the true value in the binary classification change label, that is, the unchanged label 0 or the changed label 1, represents the probability of not being predicted as the changed label or the changed label, \(W\) C represents the weight of the changed area, \(W\) nc represents the weight of the non-changed area;

[0044] The semantic change loss function \(L\) SC obtains the predicted probability distribution similarity between unchanged areas, removes the prediction of the probability distribution similarity in the changed area, and the contrast learning loss, and is specifically expressed as:

[0045]

[0046] where \(x1\) and \(x2\) are pixel vectors in the semantic segmentation results respectively, and \(y\) c is the value at the same position on \(L\) c ;

[0047] The total loss function \(L_{scd}\) is obtained by combining \(L_s\), \(L_c\) and \(L_{sc}\), and its calculation formula is as follows:

[0048]

[0049] where \(L\) s1 and \(L\) s2respectively represent the semantic segmentation losses in the dual - time images.

[0050] The present invention provides a semantic change detection system for remote sensing images, including:

[0051] An acquisition module, configured to obtain multiple pairs of remote sensing images including land cover types;

[0052] A model construction module, configured to construct an improved semantic change detection network STGNet. The improved semantic change detection network STGNet includes a dual - path dual - branch network structure and a dual - time refinement interaction structure, and input the pairs of remote sensing images into the improved semantic change detection network STGNet for training to obtain a trained semantic change detection network STGNet;

[0053] A detection module, configured to input the pairs of remote sensing images to be detected into the trained semantic change detection network STGNet to obtain the semantic change detection result of the remote sensing images;

[0054] When the model construction module inputs the pairs of remote sensing images into the improved semantic change detection network STGNet for training, it is used to obtain the features of the pairs of remote sensing images through the dual - path dual - branch network structure, and in each branch network structure of the dual - path dual - branch network structure, guide one branch to learn and fuse the feature information of the other branch; the feature information includes the context information and spatial detail information of the pairs of remote sensing images including land cover types; in the dual - time refinement interaction structure, the feature information of the two branches at different time periods is made consistent in terms of the number of channels and spatial resolution, and based on the interaction strategy, use dual - branch network structures with different scales to extract the channel attention weights in the feature information respectively, and perform feature fusion based on the channel attention weights to obtain dual - time semantic features; perform binary change detection on the dual - time semantic features to obtain a binary change detection map, and perform masking processing on the land cover types through the binary change detection map to obtain the final result of semantic change detection.

[0055] The present invention provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of the above - mentioned method for semantic change detection of remote sensing images.

[0056] The present invention provides a storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the steps of the above - mentioned method for semantic change detection of remote sensing images.

[0057] Compared with the prior art, the present invention has the following remarkable advantages:

[0058] The present invention proposes a network STGNet for guiding multi-task semantic change detection through spatio-temporal semantic interaction. Through the dual-path and dual-branch network structure of the STGNet network, cross-temporal feature-guided learning is realized for the first time in the feature extraction stage, solving the problem of missing temporal correlation caused by independent extraction of dual-temporal image features in traditional methods. It can effectively guide the network to focus more on the key change features in the dual-temporal images, and then dynamically adjust the multi-scale feature fusion strategy through channel attention weights. On the basis of unifying the channel dimension and spatial resolution, it can solve the semantic misalignment problem in the areas with blurred object boundaries, and achieve the effect of keeping the semantic regions in the semantic change detection task consistent with those detected in the segmentation task, improving the recognition ability in the semantic change detection task. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is the overall architecture diagram of STGNet in the present invention;

[0060] Figure 2 It is the diagram of the bidirectional guidance module BiDS between spatial detail information and semantics in the present invention;

[0061] Figure 3 It is the diagram of the dynamic convolution in the present invention;

[0062] Figure 4 It is the diagram of the dual-temporal refinement interaction module in the present invention;

[0063] Figure 5 It is the diagram of the feature fusion module AFF in the present invention;

[0064] Figure 6 It is the example diagram from the SECOND dataset in the present invention; where Figure 6 (a1) to Figure 6 (a5) of are the images of time period 1, Figure 6 (b1) to Figure 6 (b5) of are the images of time period 2, Figure 6 (c1) to Figure 6 (c5) of are the images of label 1, Figure 6 (d1) to Figure 6 (d5) of are the images of label 2;

[0065] Figure 7 It is the example diagram from the Landsat-SCD dataset in the present invention; where Figure 7 (a1) to Figure 7 (a5) of are the images of time period 1, Figure 7 (b1) to Figure 7 (b5) of are the images of time period 2, Figure 7 (c1) to Figure 7 (c5) of are the images of label 1,Figure 7 of (d1) to Figure 7 of (d5) is the image of label 2;

[0066] Figure 8 This is the qualitative analysis in the present invention on SECOND. The first column represents the bi-temporal remote sensing image map, the second column represents the true label, and the remaining columns represent the semantic change detection results of different methods; among them Figure 8 of (a1) to Figure 8 of (a8) is the image of time period 1 in the first group, Figure 8 of (b1) to Figure 8 of (b9) is the image of time period 2 in the first group, Figure 8 of (c1) to Figure 8 of (c9) is the image of time period 1 in the second group, Figure 8 of (d1) to Figure 8 of (d9) is the image of time period 2 in the second group, Figure 8 of (e1) to Figure 8 of (e9) is the image of time period 1 in the third group, Figure 8 of (f1) to Figure 8 of (f9) is the image of time period 2 in the third group;

[0067] Figure 9 This is the qualitative analysis in the present invention on Landsat - SCD. The first column represents the bi-temporal remote sensing image map, the second column represents the true label, and the remaining columns represent the semantic change detection result maps of different methods; among them Figure 9 of (a1) to Figure 9 of (a8) is the image of time period 1 in the first group, Figure 9 of (b1) to Figure 9 of (b9) is the image of time period 2 in the first group, Figure 9 of (c1) to Figure 9 of (c9) is the image of time period 1 in the second group, Figure 9 of (d1) to Figure 9 of (d9) is the image of time period 2 in the second group, Figure 9 of (e1) to Figure 9 of (e9) is the image of time period 1 in the third group, Figure 9 of (f1) to Figure 9 of (f9) is the image of time period 2 in the third group. Detailed implementation manner

[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] As Figures 1 - 9 shown, a semantic change detection method for remote sensing images provided by the present invention will be described below, which specifically includes the following steps:

[0070] Step S1: Obtain multiple pairs of remote sensing images containing land cover types.

[0071] Step S2: Construct an improved semantic change detection network STGNet. The improved semantic change detection network STGNet includes a dual-path dual-branch network structure and a dual-temporal refinement interaction structure, and input the dual-temporal remote sensing images into the improved semantic change detection network STGNet for training to obtain a trained semantic change detection network STGNet.

[0072] Inputting the dual-temporal remote sensing images into the improved semantic change detection network STGNet for training includes:

[0073] Step S21: Obtain the features of the dual-temporal remote sensing images through the dual-path dual-branch network structure, and in each branch network structure of the dual-path dual-branch network structure, guide one branch to learn and fuse the feature information of the other branch; the feature information includes the context information and spatial detail information of the dual-temporal remote sensing images containing land cover types.

[0074] As Figure 1 shown, the network structure is based on the common Siamese structure in CD research, constructs a dual-path feature extractor SDPNet based on the Siamese structure. This network adopts a multi-task learning architecture to separately process semantic segmentation and change detection tasks, realizing comprehensive learning of target categories and surface change information; the feature extractor includes a detail-aware path DAP and a context path CP. DDM represents a multi-task CNN branch, including a binary classification change detection task and a dual-temporal image semantic segmentation task.

[0075] The context path CP uses ResNet50 as the backbone network to extract deep semantic features, with output resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 8, and 1 / 8 of the original resolution respectively, gradually reducing the spatial resolution to enhance the feature extraction ability of semantic information. Bidirectional guidance modules BiDs are set between every two 1 / 8s; the detail-aware path DAP captures the fine spatial structure and edge information in the remote sensing image, with output resolutions of 1 / 2, 1 / 4, 1 / 4, and 1 / 4 of the original resolution respectively, and bidirectional guidance modules BiDs are set between every two 1 / 4s.

[0076] In the SCD task, the imaging range of high-resolution remote sensing images is wider, the content is rich and complex. Therefore, simple feature extraction strategies are not sufficient to obtain comprehensive and effective feature representations. Currently, using a dual-branch network structure in the encoder for feature extraction has been proven to be an effective method. This structure obtains context information and spatial detail information through different convolutional layers respectively, so as to be able to capture the key change features in the image more accurately. However, for traditional dual-branch networks, their dual-branch paths are independent, lacking mutual complementation and promotion between branches. For example, compared with the spatial detail information branch, although the context information branch is good at capturing global context, it is not sensitive enough to the subtle changes in local information. This limitation is particularly obvious in the feature extraction of high-resolution remote sensing images, because high-resolution images often contain a large amount of detail information and complex context relationships, and both global and local features need to be considered simultaneously for effective analysis and understanding.

[0077] To extract robust features with details and semantics, the idea in PIDNet was borrowed to design a module called BIDS to optimize the feature extraction ability of each branch. As Figure 1 shown in part a structure diagram of, BIDS plays a role of feature guidance in the dual-branch path, where the green arrow represents downsampling and the yellow arrow represents upsampling. The basic idea of the BIDS module is: let one branch be able to selectively learn and fuse the feature information of the other branch. Through this cross-branch feature communication, the feature representation ability of each branch is enhanced, making it more purposeful to extract features.

[0078] By designing a bidirectional guidance module BiDs between spatial detail information and deep semantic features in the dual-path extraction to learn and guide one branch to selectively learn and fuse the feature information of the other branch, thereby enhancing the feature representation ability of each path, including:

[0079] The detailed structure is as Figure 2 shown. The corresponding pixel vectors in the feature maps X and Y are respectively defined as and and dynamic convolution operations are performed. Dynamic convolution pairs and The processing procedures are as follows:

[0080]

[0081] where represents the attention weight of the k-th linear function , g is the activation function, W and b are the weight matrix and the bias vector respectively, and x represents the input feature component.

[0082] As Figure 3 shown, since the dynamic convolution endows the convolutional kernel with the attention mechanism and the generation of π k is non-linear, for different inputs, the dynamic convolution can adaptively adjust the combination mode of the convolutional kernels so that it can focus on more critical features, which has stronger feature expression ability in processing high-resolution remote sensing images.

[0083] Batch normalization is performed on these features, and element-wise multiplication and summation operations are carried out to fuse the feature information from different branches. Activation is performed through the Sigmoid function to obtain the possibility σ that two pixels belong to the same object. If σ is very high, more trust will be placed in the from its own branch, otherwise more trust will be placed in the other branch Specifically:

[0084]

[0085] where f represents the dynamic convolution operation, BN represents the batch normalization operation, Sum represents the summation operation, and Sigmod represents the activation function.

[0086] By using BIDS to achieve mutual complementarity and mutual promotion between the dual-branch paths, it is beneficial for the detail perception path to utilize the context information to enhance the understanding of local details. At the same time, the context path can also benefit from the detail information and more accurately capture the global structure. This two-way guidance between branches strengthens the extraction ability of each branch. In the processing of high-resolution remote sensing images, it is beneficial for the extractor to obtain a richer feature map, which is crucial for the subsequent SCD task.

[0087] Step S22: In the dual-temporal refinement interaction structure, keep the feature information of the dual branches at different time periods consistent in terms of the number of channels and the spatial resolution. Based on the interaction strategy, use the dual-branch network structures with different scales to respectively extract the channel attention weights in the feature information, and perform feature fusion based on the channel attention weights to obtain the dual-temporal semantic features.

[0088] To enhance the model's ability to identify unchanged regions, a Cross - Time Interaction Module (CTIM) is proposed. This module utilizes the principle of cross - learning to focus on promoting the in - depth interaction and fusion of semantic information between dual - time images, thus helping to learn richer semantic features from single - time - period images. In addition, an Attention Feature Fusion Module (AFF) is introduced, aiming to better fuse the features of semantic information and spatial information and strengthen the spatial detail information in semantic features.

[0089] Change detection is based on dual - time images. The features of remote sensing images obtained at different times are limited. Therefore

[0090] Deeply mining the internal connections and differences between dual - time images helps to identify unchanged regions in change detection. A Cross - Time Interaction Module (CTIM) is designed using a cross - learning strategy to achieve cross - time - scale feature fusion and interaction. CTIM takes the output of SDPNet as input, including low - level spatial detail features (DT1, DT2) and high - level semantic features (ST1, ST2) for time periods T1 and T2 respectively. Among them, the feature representations of the DAP and CP branches are complementary. Using simple addition or concatenation fusion methods ignores the diversity of these two types of information and will lead to performance degradation. In addition, the information in different time periods is also very different. Therefore, a hybrid aggregation layer is designed to merge information from different time periods, using the context information in the CP branch of the other time period to guide the feature response of the detail branch. First, through depth - wise separable convolution (kernel size = 3), the dual - branch features from time periods T1 and T2 are pre - processed to ensure the consistency of feature channel numbers and spatial resolutions across different time periods. Subsequently, an interaction strategy is adopted to promote the in - depth interaction and fusion of cross - time - period information.

[0091] Adopting an interaction strategy, a dual - branch network structure with different scales is used to extract channel attention weights respectively, including:

[0092] Taking time period T1 as an example, after activating the high - level semantic feature ST2 of time period T2 through the sigmoid function, it is used as a weight factor to multiply the low - level detail feature DT1 of time period T1, allowing the high - level semantic feature in time period T2 to guide the spatial detail information in time period T1, thereby obtaining cross - time - period spatial detail information FT1; The specific expression is:

[0093] F T1 =sigmod(DWConv 3×3 (Up(S T1 )))×DWConv 3×3 (D T1 )

[0094] F T2 =sigmod(DWConv3×3 (Up(S T2 )))×DWConv 3×3 (D T2 )

[0095] Where DWConv3x3 represents depthwise separable convolution with a convolution kernel size of 3x3, and UP represents upsampling.

[0096] In addition, to better fuse spatial detail features and deep semantic features, the AFF module is introduced. In the feature fusion stage, the AFF module does not adopt a simple linear fusion method, but uses a dual-branch structure with different scales to extract channel attention weights respectively, and then performs dynamic feature fusion based on the weight information. The specific structure is as Figure 4 shown. Based on the weight information, dynamic feature fusion is performed to obtain dual-temporal semantic features, including:

[0097] For the input low-level spatial detail information X and high-level semantic information Y, first perform linear addition feature fusion to obtain a preliminary fusion result F, and the calculation formula is as follows:

[0098] F = X + Y

[0099] To more accurately capture the important information in feature F, global average pooling and pointwise convolution are used to extract the global and local channel attention in feature F respectively, as Figure 5 shown. This way of processing features using different branches helps to better identify the location of the changed area and related spatial information, thereby improving the localization ability and feature expression ability of the model. Specifically, the global average pooling branch performs global average pooling on feature F to obtain a global feature vector, and then uses pointwise convolution to process this vector to extract the global channel attention weight. The pointwise convolution branch directly performs pointwise convolution on feature F to extract the local channel attention weight. The extracted attention weights are normalized to between 0 and 1 through the sigmoid function and multiplied by the corresponding features. In this way, each feature will be assigned a corresponding weight according to its importance, and the weighted features are added together to obtain the final fusion result. Through this fusion method, it is more beneficial for the model to perform dynamic adjustment during the feature fusion process, thereby improving the extraction of complex feature sets in high-resolution remote sensing images. The specific expression is:

[0100] L(F) = B(PWConv2(δ(B(PWConv1(F)))))

[0101] G(F) = B(PWConv2(δ(B(PWConv1(GAP(F))))))

[0102]

[0103] Among them, F represents the initial feature fusion result. Both PWConv1 and PWConv2 represent 1x1 pointwise convolutions, B represents the BatchNorm layer, δ represents the ReLU activation function. GAP represents the global average pooling operation.

[0104] Step S23: Perform binary change detection on the dual-temporal semantic features to obtain a binary change detection map, and perform masking processing on the land cover types through the binary change detection map to obtain the final result of semantic change detection.

[0105] Subsequently, precise binary change detection is performed by tightly combining these dual-temporal semantic features. Finally, the land cover classification result is masked by the binary change detection map, so as to accurately obtain the final result of spatial change detection (SCD), realizing high-precision detection of semantic changes in SCD tasks in remote sensing images.

[0106] The method of the present invention has three output ports, including the semantic classification result of time period 1, the semantic classification result of time period 2, and the changed part between time period 1 and time period 2. For specific understanding, refer to the overall network structure diagram. Masking processing in the present invention refers to using the changed part between time period 1 and time period 2 to filter out some parts in time period 1 and time period 2, so as to only obtain the semantic information of the changed part and get the final output. That is, the semantic change map of time period 1 and the semantic change map of time period 2.

[0107] Three loss functions are used to optimize the BCD and SS tasks in remote sensing semantic change detection, including the semantic loss function (Ls), the binary change function (Lc), and the semantic change loss function (LSC) proposed by Ding et al.

[0108] The semantic loss function Ls is designed for the semantic categories in the subtask SS, and the multi-class cross-loss entropy is calculated for the semantic classification results S1 and S2 in the SS task and the semantic categories TS1 and TS2 in the true semantic change map:

[0109]

[0110] where y i and represent the true semantic label category and the probability predicted as the i-th category respectively. N represents the land cover types in semantic change detection, and 0 as the unchanged category is excluded because this is more conducive to the model focusing on the extraction of semantic features in the changed areas.

[0111] The binary change function Lc is the weighted binary cross-entropy (WBCE) loss between the predicted change map and the ground truth change map in the BCD task, which is used to balance the class imbalance between the changed and unchanged regions in the BCD task. The calculation formula is as follows:

[0112]

[0113] where y c represents the ground truth value in the binary change label, i.e., the unchanged label 0 or the changed label 1, represents the probability of not being predicted as the changed label or the changed label. N represents the number of image pixels, WC represents the weight of the changed region, and Wnc represents the weight of the unchanged region, which are set to 0.25 and 0.75 respectively.

[0114] Lsc is a contrastive learning-based loss function that connects the BCD task and the SS task. Specifically, in the overall task of semantic change detection, the Lsc loss function encourages the prediction of similar probability distributions between unchanged regions, but penalizes the prediction of similar probability distributions in changed regions. The calculation formula is as follows

[0115]

[0116] where x1 and x2 are pixel vectors in the semantic segmentation results respectively, and y c is the value at the same position as Lc.

[0117] Thus, the total loss function Lscd is obtained by combining Ls, Lc, and Lsc, and its calculation formula is as follows:

[0118]

[0119] where Ls1 and Ls2 represent the semantic segmentation losses in the dual-time images respectively.

[0120] Step S3: Input the dual-time remote sensing images to be detected into the trained semantic change detection network STGNet to obtain the semantic change detection results of the remote sensing images.

[0121] Dataset: To better verify the proposed network, experiments are conducted on two publicly available semantic datasets: the SECOND dataset and the Landsat-SCD dataset.

[0122] The SECOND dataset consists of 4,662 pairs of remotely sensed images collected from multiple platforms and sensors, covering multiple important urban areas in China, including Hangzhou, Chengdu, Shanghai, etc., and encompassing a rich variety of land cover types. Each image in this dataset has a size of 512×512 pixels, with a spatial resolution ranging from 0.5 meters to 3 meters, capable of capturing subtle surface changes. However, currently, only 2,968 pairs of dual-temporal image pairs with true labels are available, and the changing pixels account for 19.87% of the total pixels in the images. As Figure 6 shown in the example, this dataset defines 7 categories of labels, including a no-change category and five land cover change categories (bare ground, trees, low vegetation, water, buildings, and playgrounds). To scientifically and reasonably evaluate the model performance, the SECOND dataset is divided into a training set and a test set in a ratio of 4:1.

[0123] The images in the Landsat-SCD dataset contain three basic bands: red (R), green (G), and blue (B), and have a spatial resolution of 30 meters. As Figure 7 shown in the example, this dataset defines five label categories, including a no-change category and four land cover change categories (farmland, desert, buildings, and water bodies). The changing pixels account for 18.89% of the total pixels in the images, providing rich surface change information. The Landsat SCD dataset contains 8,468 pairs of images, and each image has a size of 416×416 pixels. Excluding image enhancement, this data has 2,385 original image pairs, which are divided into a training set, a validation set, and a test set in a ratio of 3:1:1.

[0124] In the study, to better quantitatively analyze the experimental results, 4 accuracy metrics commonly used in the SCD task were adopted to evaluate the proposed method, namely overall accuracy (OA), mean intersection over union (mIOU), separation kappa (SeK), and F1scd. The OA metric measures the proportion of pixels correctly classified in all categories to the total pixels, and can provide a global and intuitive accuracy assessment. Assume Q = {q i,j} is the confusion matrix, q i,j represents the number of pixels classified as class i, and j represents the number of pixels in the true class (i, j ∈ {0, 1,..., N}, 0 represents no change). Then the calculation formula for OA is as follows:

[0125]

[0126] In the semantic change detection dataset, the unchanged part occupies the majority, and the OA metric is determined by identifying unchanged pixels, which cannot accurately identify the land use / land cover (LULC) types. Therefore, two additional metrics, mIoU and SeK, are adopted to evaluate the segmentation performance of the model for the changed and unchanged regions in the BCD task, and the classification accuracy for different LULC types in the SS task, respectively.

[0127] mIoU is a metric commonly used to evaluate image segmentation performance. It measures the accuracy of segmentation by calculating the ratio of the intersection to the union between the ground truth segmentation region and the predicted segmentation region. In the BCD task, mIoU consists of the unchanged region (IoUnc) and the changed region (IoUc), and the calculation formula is as follows:

[0128] mIoU = (IoU nc + IoU c ) / 2

[0129]

[0130] Sek is used to evaluate the semantic categories of the transformed regions. It is calculated based on the confusion matrix where let q 00 = 0 to exclude the truly unchanged pixels that dominate in quantity. The calculation formula is as follows:

[0131]

[0132] F1scd was proposed by Ding et al. to evaluate the segmentation accuracy of LULC classes in the changed regions. This metric calculates the precision (Pscd) and recall (Rscd) of the changed regions by referring to F1score. The calculation formula is as follows:

[0133]

[0134] Experimental settings: All experiments were implemented on an NVIDIA GPU (GeForce RTX 4060Ti) with 16GB of video memory. In all experiments, the same experimental parameter configurations were adopted, specifically including: setting the batch size to 8, conducting 80 rounds of training, and setting the initial learning rate to 0.1. In addition, the Stochastic Gradient Descent (SGD) optimizer was used to iteratively update the model parameters to minimize the loss function.

[0135] Experimental comparison and analysis:

[0136] Comparison method: To better compare the superiority of the method of the present invention in identifying changed areas and land cover types in the SCD task, comparative analysis is carried out with the six latest existing methods, including HRSCD3, HRSCD4, BiSRNet, SCanNet, HGINet, and STSP-Net. To ensure fair comparative experiments, no pre-trained weights are used in the training, and no data augmentation is performed.

[0137] HRSCD3: By fusing multi-scale features and constructing a BCD branch to introduce time-related information into the network, the detection of land cover type changes is achieved.

[0138] HRSCD4: As an upgraded version of the HRSCD3 series, HRSCD4 performs skip connections while maintaining the ability of high-resolution semantic change detection, connecting the siamese encoder with the decoder of the CD branch to improve the recognition ability for complex change scenarios.

[0139] BiSRNet: It is a dual-temporal semantic reasoning network. By introducing a cross-temporal SR (Cot-SR) block to model the time correlation, the recognition ability for changed areas and land cover types is improved.

[0140] SCanNet: It is a multi-task-based semantic change detection network. This method constructs a semantic change transformer (SCanFormer) to explicitly construct the "from-to" semantic transformation between dual-temporal RSIs, achieving effective recognition of changed areas and land cover types in remote sensing images.

[0141] HGINet: It is a semantic change detection network based on hierarchical semantic graph interaction. By modeling the dual-temporal correlation and using graph learning to represent the interaction of different feature layers, changed areas and land cover types are accurately recognized.

[0142] STSP-Net: It is a spatio-temporal semantic perception network. By introducing a spatio-temporal attention mechanism, spatio-temporal information in the image is effectively captured, thus achieving the recognition of changed areas and land cover types.

[0143] Quantitative and qualitative analysis:

[0144] On the Second dataset: Table 1 shows the quantitative comparison results of the method with other methods on the Second dataset. Compared with other methods, the proposed STGNet performs significantly better than other methods on the Second dataset, ranking first in all indicators. Specifically, the mIoU reaches 72.83%, Sek is 22.45%, F1scd is 61.83%, and OA is 87.51%. In the comparison of these methods, HRSCD3 is a network that directly acts on dual-temporal images for binary change detection, while other methods mainly perform change detection based on the feature map information obtained from dual-temporal images. Therefore, HRSCD3 has the worst mIoU value on the high-resolution Second dataset, only achieving 66.85%, while its upgraded version, HRSCD4, which performs CD change detection by extracting features through an encoder, has a significant improvement in mIoU, reaching 72.08%. This proves the importance of first extracting features and then performing change detection in the SCD task. The proposed SDPNet extractor and BIDS module are designed to better extract features in high-resolution remote sensing images and deeply fuse dual-temporal image features through CTIM. In the ablation experiment, its specific advantages will be introduced in detail. Therefore, for STSP-Net, which also focuses on the feature extraction stage, its mIoU reaches the second, at 72.31%. To visually demonstrate the advantages of the method on the Second dataset, 3 pairs of dual-temporal remote sensing images are selected for qualitative analysis, and the details are highlighted with red rectangular frames. As Figure 8 shown in the third and fourth columns of

[0145] Table 1: Quantitative comparison of different methods on the Second dataset, with the best results shown in bold black.

[0146] Method mIoU (%) Sek (%) F1scd (%) OA (%) HRSCD3 66.85 12.99 52.99 84.99 HRSCD4 72.08 20.3 59.67 86.76 BiSRNet 69.61 17.13 57.04 85.65 SCanNet 71.56 20.04 59.80 86.86 HGINet 67.94 13.55 52.91 84.89 STSP - Net 72.31 20.25 59.37 86.84 Our 72.83 22.45 61.83 87.51

[0147] On the Landsat-SCD dataset: To further study the robustness and performance of the proposed method, quantitative and qualitative analyses were also carried out on the lower-resolution Landsat-SCD dataset. The experimental results show that the new method also has significant advantages when facing low-resolution data. As shown in Table 2, compared with STSP-Net, which ranks second on this dataset, STGNet has significant improvements in all evaluation indicators. Specifically, the mIoU has increased by 5.83%, Sek has increased by 16.66%, F1scd has increased by 6.56%, and the OA accuracy has increased by 2.30%. These data fully prove the superior performance and strong robustness of the method in the low-resolution remote sensing image classification task.

[0148] Table 2: Quantitative comparison of different methods on the Landsat-SCD dataset

[0149] Method mIoU (%) Sek (%) F1scd (%) OA (%) HRSCD3 79.79 35.57 75.86 91.47 HRSCD4 81.07 38.09 77.37 92.17 BiSRNet 69.99 17.20 62.79 87.93 SCanNet 84.82 49.26 84.38 94.37 HGINet 76.20 28.56 71.67 90.07 STSP - Net 85.10 49.91 84.72 94.52 Our 90.93 66.57 91.28 96.82

[0150] To more intuitively show the performance of the method on low-resolution images, partial visualization results are provided in Figure 9 By comparison, it can be found that when other methods identify the changed areas of low-resolution images, many key detail contour information is often lost. For example, the BiSRNet model, which performs well in the Second dataset, encounters obvious challenges when applied to the Landsat-SCD dataset for change detection, losing a large amount of detail contour information. While in the dual-branch, BiDS is used to enhance the extraction ability of the image. Therefore, even in low-resolution images, good performance can still be achieved.

[0151] The ablation experiments are as follows:

[0152] To further verify the effectiveness of each part of the method proposed in the present invention, ablation verification experiments were comprehensively implemented on the Landsat-SCD dataset, specifically covering the following key aspects: the selection of the backbone network, the optimization of the dual-branch feature extraction path, the enhancement effect of the BiDS module on feature extraction in the dual-branch feature extraction, the performance of the CTIM module as a dual-temporal interaction for dual-temporal remote sensing semantic detection, and the computational cost of the model.

[0153] As shown in Table 3, using the SSCDL architecture proposed by Lei et al. as the basic model, it can be seen that with a small increase in the number of parameters and computational volume, SSCDL using Resnet50 has significantly improved all indicators compared to SSCDL using Resnet34. This is because Resnet50 has a deeper network structure and stronger feature extraction ability, and can learn more complex and rich feature representations, which is suitable for processing complex and variable image data. In addition, by adding the spatial detail branch information DAP, it can be found that all its indicators have been significantly improved. Compared with the basic network based on Resnet50, it has an mIoU improvement of 4.63%, a Sek improvement of 12.23%, an F1scd improvement of 5.59%, and an OA improvement of 2.06%. This fully shows that the introduction of the spatial detail information branch can effectively improve the performance of the model and enhance the ability to capture spatial detail information, thus achieving significant improvements in multiple evaluation dimensions. BiDS plays a two-way guiding role in the dual-branch structure. Therefore, after adding BiDS, the network pays more attention to the extraction of important information in remote sensing images.

[0154] Therefore, it further improved by 1.3%, 3.91%, 1.31%, and 0.45% in terms of mIoU, Sek, Fscd, and OA respectively. With the addition of CTIM, the proposed method achieved the best results in all metrics on the Landsat-SCD dataset. Specifically, the mIoU reached 90.93%, Sek was 66.57%, FSCD was 91.28%, and OA was 96.82%. This fully verified that the proposed CTIM can deeply explore the internal connections and differences between dual-temporal images, thus contributing to the identification of unchanged areas in change detection. It should be noted that the addition of BIDS and CTIM only increased the number of parameters by 1% and 0.03% respectively. It can be seen that better change detection and land category identification were achieved with the introduction of a small number of parameters.

[0155] Table 3: Ablation experiment results on the Landsat-SCD dataset

[0156]

[0157] In the SCD (Semantic Change Detection) task, it is crucial to accurately identify the change regions and change types. In this study, several common challenges in existing SCD tasks were deeply analyzed, and a network (i.e., STGNet) for semantic change detection by guiding multi-tasks through spatio-temporal semantic interaction was proposed. The network adopted a dual-path feature extractor based on the Siamese structure, and by cleverly integrating spatial detail information, significantly enhanced the ability to extract complex feature sets of remote sensing images. On this basis, an innovative bidirectional guidance module (BiDS) was further designed, which can establish an effective connection between spatial detail information and high-level semantics, further strengthening the ability of the feature extractor to capture key information and enriching the representation of deep semantic features and low-level spatial information features.

[0158] In addition, to make full use of the temporal correlation between dual-temporal images (i.e., images at different time points), a dual-temporal refinement interaction module (CTIM) was carefully designed. This module deeply explores the internal connections and subtle differences between dual-temporal images, thus contributing to more accurately identifying the unchanged areas in change detection and improving the overall change detection accuracy.

[0159] To comprehensively evaluate the performance of the proposed network, exhaustive experiments were conducted on two publicly available authoritative datasets. The experimental results show that compared with the state-of-the-art change detection methods, the proposed STGNet achieves the highest accuracy with a small number of parameters introduced, which further verifies the effectiveness of the network in dealing with complex remote sensing image change detection tasks. In future work, the focus will continue to be on how to achieve semi-supervised semantic change detection to further reduce the over-reliance on datasets, lower the cost of manual annotation, and improve the generalization ability of the model under limited annotated data. Specifically, a large amount of unannotated data will be combined with a small amount of high-quality annotated data to deeply explore effective semi-supervised learning algorithms. These algorithms will be dedicated to accurately capturing subtle changes at the semantic level while maintaining the efficiency and practicality of the model, providing more powerful technical support for the semantic change detection task of remote sensing images.

[0160] Based on the above method, the present invention provides a semantic change detection system for remote sensing images, including: an acquisition module, a model construction module, and a detection module.

[0161] Among them, the acquisition module is used to obtain multiple pairs of remote sensing images containing land cover types; the model construction module is used to construct an improved semantic change detection network STGNet. The improved semantic change detection network STGNet includes a dual-path dual-branch network structure and a dual-temporal refinement interaction structure, and inputs the dual-temporal remote sensing images into the improved semantic change detection network STGNet for training to obtain a trained semantic change detection network STGNet; the detection module is used to input the dual-temporal remote sensing images to be detected into the trained semantic change detection network STGNet to obtain the semantic change detection result of the remote sensing images.

[0162] When the model construction module inputs the dual-temporal remote sensing images into the improved semantic change detection network STGNet for training, it is used to obtain the features of the dual-temporal remote sensing images through the dual-path dual-branch network structure, and in each branch network structure of the dual-path dual-branch network structure, guide one branch to learn and fuse the feature information of the other branch; the feature information includes the context information and spatial detail information of the dual-temporal remote sensing images containing land cover types; in the dual-temporal refinement interaction structure, the feature information of the dual branches at different time periods is made consistent in terms of the number of channels and spatial resolution, and based on the interaction strategy, dual-branch network structures with different scales are used to extract the channel attention weights in the feature information respectively, and feature fusion is performed based on the channel attention weights to obtain dual-temporal semantic features; binary change detection is performed on the dual-temporal semantic features to obtain a binary change detection map, and the land cover types are masked by the binary change detection map to obtain the final result of semantic change detection.

[0163] The present invention also provides a computer device, including a memory and a processor. A program is stored in the memory. When the program is executed by the processor, the processor is caused to execute the steps of a semantic change detection method for remote sensing images.

[0164] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.

[0165] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a semantic change detection method for remote sensing images are implemented.

[0166] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0167] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A semantic change detection method for remote sensing images, characterized in that, Including: Obtain multiple pairs of remote sensing images including land cover types; Construct an improved semantic change detection network STGNet, where the improved semantic change detection network STGNet includes a dual-path dual-branch network structure and a dual-temporal refinement interaction structure, and input the pairs of remote sensing images into the improved semantic change detection network STGNet for training to obtain a trained semantic change detection network STGNet; Input the pairs of remote sensing images to be detected into the trained semantic change detection network STGNet to obtain the semantic change detection result of the remote sensing images; Among them, inputting the pairs of remote sensing images into the improved semantic change detection network STGNet for training includes: Obtain the features of the pairs of remote sensing images through the dual-path dual-branch network structure, and in each branch network structure of the dual-path dual-branch network structure, guide one branch to learn and fuse the feature information of the other branch; the feature information includes the context information and spatial detail information of the pairs of remote sensing images including land cover types; In the dual-temporal refinement interaction structure, keep the feature information of the two branches at different time periods consistent in terms of the number of channels and spatial resolution, and based on the interaction strategy, use dual-branch network structures with different scales to extract the channel attention weights in the feature information respectively, and perform feature fusion based on the channel attention weights to obtain dual-temporal semantic features; Perform binary change detection on the dual-temporal semantic features to obtain a binary change detection map, and perform masking processing on the land cover types through the binary change detection map to obtain the final result of semantic change detection.

2. The semantic change detection method for remote sensing images according to claim 1, characterized in that, The dual-path dual-branch network structure includes a detail-aware path DAP and a context path CP. The context path CP uses ResNet50 as the backbone network to extract semantic features, and the output resolutions are 1 / 2, 1 / 4, 1 / 8, 1 / 8, and 1 / 8 of the original resolution respectively. Bidirectional guiding modules BiDs are set between the two 1 / 8 resolutions; The detail-aware path DAP captures the fine spatial structure and edge information in the remote sensing image, and the output resolutions are 1 / 2, 1 / 4, 1 / 4, and 1 / 4 of the original resolution respectively. Bidirectional guiding modules BiDs are set between the two 1 / 4 resolutions.

3. A semantic change detection method for remote sensing images according to claim 1, characterized in that, The guiding one branch to learn and fuse the feature information of the other branch includes: Define the corresponding pixel vectors in feature maps A and B as and respectively, and perform dynamic convolution operations on the corresponding pixel vectors. The processing procedures for and are as follows: where, π k represents the attention weight of the k-th linear function , g is the activation function, W and b are the weight matrix and the bias vector respectively, and x represents the input feature component; Perform batch normalization processing, element-wise multiplication, and summation operations on the features, fuse the feature information from different branches, and process the feature information through the Sigmoid function to obtain the possibility σ that two pixels belong to the same object. If σ is higher than the threshold, trust the one from its own branch Specifically: Among them, f represents the dynamic convolution operation, BN represents the batch normalization operation, Sum represents the summation operation, and Sigmod represents the activation function.

4. The semantic change detection method for remote sensing images according to claim 1, characterized in that The based on the interaction strategy, using dual-branch network structures with different scales to extract the channel attention weights in the feature information respectively includes: Based on the T1 period, the high-level semantic features ST2 of the T2 period are activated through the sigmoid function, and the activated semantic features ST2 are multiplied by the detail features DT1 of the T1 period as a weight factor, prompting the spatial detail information in the T1 period to learn the semantic features in the T2 period and obtaining cross-period spatial detail information F T1 ; The specific expression is: F T1 = sigmod(DWConv 3×3 (Up(S T1 ))) × DWConv 3×3 (D T1 ); Among them, DWConv 3x3 represents depthwise separable convolution with a kernel size of 3x3, and Up represents upsampling.

5. A semantic change detection method for remote sensing images according to claim 1, characterized in that, The based on the channel attention weights to perform feature fusion to obtain dual-temporal semantic features includes: Perform linear addition feature fusion on the input spatial detail information X and semantic information Y to obtain a preliminary fusion result F, and the specific expression is as follows: F = X + Y Global average pooling and pointwise convolution are used to extract global channel attention and local channel attention in F respectively. The extracted attention weights are normalized to between 0 and 1 through the sigmoid function, and the attention weights are multiplied and added to the corresponding features to obtain the final fusion result. The specific expression is as follows: L(F) = B(PWConv2(δ(B(PWConv1(F))))); G(F) = B(PWConv2(δ(B(PWConv1(GAP(F)))))); where F represents the initial feature fusion result, PWConv1 and PWConv2 both represent 1x1 pointwise convolutions, B represents the BatchNorm layer, δ represents the ReLU activation function, and GAP represents the global average pooling operation.

6. The semantic change detection method for remote sensing images according to claim 1, characterized in that, When training the dual-temporal remote sensing images by inputting them into the improved semantic change detection network STGNet, it further includes: Adopt the semantic loss function L s , binary change function L c and semantic change loss function L SC Optimize the semantic change detection results to obtain the optimized semantic change detection network STGNet; specifically: The semantic loss function L s Performs multi-class cross-entropy loss calculation on the semantic classification results S1 and S2 and the semantic categories TS1 and TS2 in the true value semantic change graph, specifically expressed as: where y i and represent the true semantic label category and the probability of being predicted as the i-th category, respectively, and N represents the land cover type in semantic change detection; The binary variation function L c Obtain the binary cross-entropy loss between the predicted change map and the true change map, which is specifically expressed as: where y c represents the true value in the binary classification change label, that is, the unchanged label 0 or the changed label 1, represents the probability of not being predicted as the changed label or the changed label, W C represents the weight of the changed area, W nc represents the weight for the changed area; The semantic change loss function L SC Obtain the prediction of the similarity probability distribution between unchanged regions, remove the prediction of the similarity probability distribution in the changed regions, and the contrastive learning loss, which is specifically expressed as: where x1 and x2 are pixel vectors in the semantic segmentation result, and y c is the value at the same position on L c ; The total loss function Lscd is obtained by combining Ls, Lc, and Lsc, and its calculation formula is as follows: Among them, L s1 and L s2 respectively represent the semantic segmentation losses in the dual-time images.

7. A semantic change detection system for remote sensing images, characterized in that, It includes: An acquisition module for acquiring multiple dual-temporal remote sensing images containing land cover types; A model construction module for constructing an improved semantic change detection network STGNet. The improved semantic change detection network STGNet includes a dual-path dual-branch network structure and a dual-temporal refinement interaction structure, and inputs the dual-temporal remote sensing images into the improved semantic change detection network STGNet for training to obtain a trained semantic change detection network STGNet; A detection module for inputting the dual-temporal remote sensing images to be detected into the trained semantic change detection network STGNet to obtain the semantic change detection result of the remote sensing images; When the model construction module inputs the dual-temporal remote sensing images into the improved semantic change detection network STGNet for training, it is used to obtain the features of the dual-temporal remote sensing images through the dual-path dual-branch network structure, and in each branch network structure of the dual-path dual-branch network structure, guide one branch to learn and fuse the feature information of the other branch; The feature information includes the context information and spatial detail information of the dual-temporal remote sensing images containing land cover types; in the dual-temporal refinement interaction structure, the feature information of the dual branches in different time periods is made consistent in terms of the number of channels and spatial resolution, and based on the interaction strategy, the dual-branch network structures with different scales are used to extract the channel attention weights in the feature information respectively, and feature fusion is performed based on the channel attention weights to obtain the dual-temporal semantic features; Perform binary change detection on the dual-temporal semantic features to obtain a binary change detection map, and perform masking processing on the land cover types through the binary change detection map to obtain the final result of semantic change detection.

8. A computer device, characterized in that, It includes a memory and a processor. A program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a semantic change detection method for a remote sensing image according to any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of a semantic change detection method for a remote sensing image according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal remote sensing data change detection method

    CN117523401A

  • Semantically-aware image-based visual localization

    US20200357143A1

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1

  • Joint 3D detection and segmentation using bird's eye view and perspective view

    US20250054286A1

  • AU2021106836A4

Cited By

  • Satellite image in-orbit change detection method and device, storage medium and electronic equipment

    CN121438011A

  • Satellite image on-orbit change detection method and device, storage medium and electronic equipment

    CN121438011B