A Semantic Change Detection Method Based on High-Resolution Convolutional Network and Context Information Encoding

Through the high-resolution convolutional network and context information encoding method, the problem of low semantic change detection accuracy is solved, the accurate identification of substantial changes and the exclusion of non-substantial changes are achieved, and the detection accuracy is improved.

CN116310811BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310203630.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-06
Publication Date
2025-07-11
Estimated Expiration
2043-03-06

AI Technical Summary

Technical Problem

The results of existing semantic change detection methods are low in accuracy and are easily affected by non-substantial factors such as light changes, shadows and seasonal vegetation changes, resulting in low detection accuracy.

Method used

Using a method based on high-resolution convolutional network and context information encoding, multi-level and multi-scale feature information is extracted through twin high-resolution feature extraction modules and context information encoding modules, combining differential feature extraction and semantic segmentation, the network is trained using a deep supervision method to exclude the impact of non-substantial changes.

Benefits of technology

The detection accuracy of semantic change areas is improved, substantial changes can be accurately identified, and the influence of interference factors such as light and shadow is eliminated. The detection results reach 39.31% comprehensive index Score on the SECOND data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310811B_ABST
    Figure CN116310811B_ABST
Patent Text Reader

Abstract

The present invention relates to a high-resolution convolutional network and a semantic change detection method for context information encoding. A high-precision semantic change detection network model is adopted. The model includes a twin high-resolution feature extraction module and a context information encoding module. The twin high-resolution feature extraction module is used to extract the feature information of the original image pair; the context information encoding module is used for change detection and semantic segmentation: 1) Context information encoding is performed on the difference information between the feature information of the original image pair to obtain a change binary map with the same size as the input image; 2) Context information encoding is performed on the respective feature information of the original image pair to obtain two semantic segmentation maps with the same size as the input image. Finally, the change binary map and the two semantic segmentation maps are combined to obtain a semantic change detection result map. The details of the semantic change region obtained by the method proposed by the present invention are more accurate, and the accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and particularly relates to a semantic change detection method based on a high-resolution convolutional network and context information encoding. Background Art

[0002] Semantic change detection is a pixel-level change detection with more information. It not only provides a binary classification change result map, but also provides a "from-to" change result map indicating the change direction. "From-to" refers to the change in land cover types between two temporal remote sensing images, such as "from land to building", "from forest to farmland", etc. Its goal is to detect the changed areas in a pair of remote sensing images and the semantic types to which each pixel in the changed areas belongs. Compared with binary classification changes, semantic change detection is a more complex change detection task and can obtain comprehensive change information. This technology plays an important role in the fields of environmental monitoring, urban coverage, resource management, etc.

[0003] In recent years, thanks to the rapid development of high-spatial-resolution and multi-temporal remote sensing Earth observation, a large number of bi-temporal and high-spatial-resolution remote sensing images can be obtained, which provides a reliable data source for semantic change detection. Semantic change detection can simultaneously determine the areas where changes occur and the types of changes that occur. Its results consist of two separate classification maps, with different colors representing the unchanged and changed classes for each image cycle. Generally speaking, in the semantic change detection result map, white represents the areas where no changes have occurred, and other colors represent the types of land cover changes. Deep learning simplifies the traditional non-end-to-end detection methods through an end-to-end detection method, effectively improving the detection efficiency and accuracy. Therefore, semantic change detection methods based on deep learning have obtained rapid development. For example, Mou et al. proposed a new recurrent convolutional neural network architecture in the literature "Learning Spectral-Spatial-Temporal Features via a Recurrent Convolutional Neural Network for Change Detection in Multispectral Imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2019.", which combines a convolutional neural network and a recurrent neural network into an end-to-end network. The former can generate rich spectral-spatial feature representations, while the latter effectively analyzes the temporal correlation in bi-temporal images. For example, Yang et al. proposed an asymmetric siamese network ASN (Asymmetric Siamese Network) for semantic change detection in the literature ". Asymmetric siamese networks for semantic change detection in aerial images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60:1-18.", which locates and identifies semantic changes through feature pairs obtained from modules with widely different structures. These feature pairs involve different spatial ranges and numbers of parameters to consider the differences in different land cover distributions. And a semantic change detection dataset named SECOND (SEmantic Change detectiON Dataset) was further proposed and evaluated.

[0004] At present, there are many studies on change detection, but there are few exploratory studies on semantic change detection, and the overall accuracy of the results of semantic change detection is not high, and the details are poor. In addition, there are also some problems of non-substantive change detection, such as lighting changes, seasonal vegetation changes, building shadow coverage, etc. These all belong to the significant differences in visual appearance between two-phase images, that is, non-substantive changes, not the desired substantive changes. These interference factors will affect the accuracy of the detection results. Therefore, it is very necessary to design a high-precision semantic change detection network. Summary of the Invention

[0005] Technical Problems to be Solved

[0006] Aiming at the problem of the low accuracy of the results of existing semantic change detection methods, the present invention provides a semantic change detection method based on a high-resolution convolutional network and context information encoding.

[0007] Technical Solution

[0008] A semantic change detection method based on a high-resolution convolutional network and context information encoding, characterized in that the steps are as follows:

[0009] S1. Respectively input remote sensing images T1 and T2 of different time phases into a siamese high-resolution feature extraction module after changing the number of channels through convolutional modules C0 and C1, and perform feature extraction and exchange of feature information of different scales through multiple small convolutional modules in the feature extraction module to obtain feature maps F 1 j ’ and F 2 j ’; Unify the resolution and the number of channels of the features F 1 j ’ and F 2 j ’ through convolution and upsampling operations respectively to obtain new feature information F 1 j and F 2 j ;

[0010] S2. Perform differential feature extraction on the feature pair F 1 j and F 2 j to obtain the differential information feature d j of the two images, stack the d j channels and perform context information encoding, and obtain a change binary map O with the same size as the input image after post-processing;

[0011] S3. The feature pair F 1 j and F2 j Perform channel stacking separately, and then perform context information encoding respectively to obtain two semantic segmentation maps S1 and S2 with the same size as the input image, corresponding to the original images T1 and T2 respectively;

[0012] S4, combine the semantic segmentation maps S1 and S2 with the change detection map O to obtain the final semantic change detection result maps O1 and O2.

[0013] A further technical solution of the present invention: In step S1, the siamese high-resolution feature extraction module includes two feature extraction branches, and the weights are shared between the two branches.

[0014] A further technical solution of the present invention: The high-resolution network model includes a plurality of small convolution modules. All the convolution modules in the upper and lower high-resolution network model branches are respectively named convolution module C 1 i,j and convolution module C 2 i,j , i≥1, j≥0. The feature map passes through convolution module C 1 i,j and convolution module C 2 i,j to obtain a new feature map I 1 i,j and I 2 i,j . The resolution of feature map I 1 i,j and I 2 i,j is denoted as H 1 i,j ×W 1 i,j and H 2 i,j ×W 2 i,j , and the number of channels is denoted as C 1 i,j and C 2 i,j , where H 1 i,j =H 2 i,j =H input / 2 j , W 1 i,j =W 2 i,j =W input / 2 j , C 1 i,j =C 2 i,j =32×2j , where H input and W input are the resolution sizes of the input image pair T.

[0015] A further technical solution of the present invention: In step S1, the input of the convolutional modules C 1 i,j and C 1 i,j comes from the outputs of the convolutional modules C 1 i-1,y and C 2 i-1,y , where i≥2, y∈[0, i - 2]. Here, the output of C 1 i-1,y is denoted as I 1 i-1,y , and the output of C 2 i-1,y is denoted as I 2 i-1,y ; separately, the input of the convolutional modules C 1 1,0 and C 2 1,0 comes from the feature maps obtained by T1 and T2 through the convolutional module C0 and the convolutional module C1.

[0016] A further technical solution of the present invention: In step S1, the convolutional modules C 1 i,j and C 2 i,j have multiple inputs with different resolutions and numbers of channels, and it is necessary to unify the resolutions and numbers of channels for addition and fusion; the rules for changing the resolutions and numbers of channels are as follows: For the input I i-1,y (i≥2, y∈[0, i - 2]), when y < j, perform (j - y) strided convolutions with a stride of 2 on the feature map I i-1,y . Each strided convolution doubles the number of channels of the feature map and halves the resolution through a 3×3 convolution; when y = j, perform a 3×3 convolution on the feature map I i-1,y for feature extraction with the number of channels and resolution unchanged; when y > j, perform a 3×3 convolution on the feature map I i-1,y , change the number of channels to 32×2 j , and use bilinear interpolation upsampling operation to make the resolution become H input / 2 j ×W input / 2 j . After unifying the resolutions and numbers of channels of the multiple feature map inputs of the convolutional modules C 1 i,j and C 2 i,j , fuse the feature maps by addition.

[0017] A further technical solution of the present invention: In step S2, for the feature pair F 1 j and F 2 j performing differential feature extraction means taking the absolute difference between the corresponding feature pairs F 1 j and F 2 j as the differential feature d i .

[0018] A further technical solution of the present invention: The post-processing in step S2 refers to performing binarization processing using threshold segmentation to obtain the final detection result map O.

[0019] A further technical solution of the present invention: Combining the semantic segmentation maps S1 and S2 with the change detection map O in step S4 means only retaining the semantic segmentation results of the regions in the semantic segmentation maps S1 and S2 that are the same as the changed regions in the change detection map O, and ignoring the semantic segmentation results of the regions in O that have not changed.

[0020] A further technical solution of the present invention: After step S4, it further includes the step of guiding the prediction network using the deep supervision method. When training the network using the deep supervision method, obtaining the binary cross-entropy loss Lbce from the change binary map O and the true change label, obtaining the cross-entropy losses L1ce and L2ce from the semantic change detection result maps O1 and O2 and the true semantic change labels respectively; adding L1ce, L2ce, and Lbce together with weights to obtain the total loss L, and performing backpropagation on L, repeating the iteration until the iteration number reaches the set initial value to determine that the training is completed.

[0021] A further technical solution of the present invention: All convolutional modules are composed of a 3×3 convolutional layer, a batch normalization layer, and a rectified linear unit.

[0022] Beneficial effects

[0023] A semantic change detection method based on a high-resolution convolutional network and context information encoding provided by the present invention has more accurate details in the obtained semantic change regions, and the accuracy is effectively improved; it can exclude the influence of non-substantive changes caused by interference factors such as illumination, shadow, and seasonal changes. In terms of accuracy, the OCHRSCD of the present invention reaches a comprehensive index Score of 39.31% on the SECOND dataset, and the details of the detected semantic change regions are more accurate.

[0024] 1. Using a high-resolution convolutional network to extract features, the high-resolution network can extract multi-level and multi-scale feature information and simultaneously retain high-resolution fine features;

[0025] 2. Context information encoding is adopted. The following information encoding module can enhance the correlation between pixels and the target area and strengthen the regional connection of similar pixels.

[0026] 3. The details of the semantic change area obtained by combining the high-resolution convolutional network and context information encoding are more accurate, and the accuracy is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0028] Figure 1 It is a network structure diagram of the method of the embodiment of the present invention.

[0029] Figure 2 It is a structure diagram of the context information encoding module in the network model of the embodiment of the present invention.

[0030] Figure 3 It is a comparison table of test results of the method of the embodiment of the present invention and other existing methods.

[0031] Figure 4 It is a schematic diagram of the semantic change detection result of the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0033] The present invention designs a semantic change detection method based on a high-resolution convolutional network and context information encoding. By constructing a new remote sensing image semantic change detection model based on a high-resolution network and context information encoding, it is used for semantic change detection of high-resolution remote sensing images. As Figure 1 shown, the detection model includes two parts: a twin high-resolution feature extraction module and a context information encoding module. The twin high-resolution feature extraction module is used to extract the feature information of the original image pair; the context information encoding module is used for change detection and semantic segmentation: 1) Context information encoding is performed on the difference information between the feature information of the original image pair to obtain a change binary map with the same size as the input image; 2) Context information encoding is performed on the respective feature information of the original image pair to obtain two semantic segmentation maps with the same size as the input image. Finally, the semantic change detection result map is obtained by combining the change binary map and the two semantic segmentation maps. The context information encoding module is as Figure 2as shown

[0034] The specific method includes the following steps:

[0035] S1. Respectively input remote sensing images T1 and T2 in different phases into the siamese high-resolution feature extraction module after changing the number of channels through convolution modules C0 and C1. Perform feature extraction and exchange of feature information at different scales through multiple small convolution modules in the feature extraction module to obtain feature maps F 1 j ’ and F 2 j ’; Unify the resolution and the number of channels of the features F 1 j ’ and F 2 j ’ through convolution and upsampling operations respectively to obtain new feature information F 1 j and F 2 j ;

[0036] S2. Perform differential feature extraction on the feature pair F 1 j and F 2 j to obtain the differential information feature d j of the two images. Stack the d j channels and perform context information encoding. After post-processing, obtain a binary change map O with the same size as the input image;

[0037] S3. Stack the channels of the feature pair F 1 j and F 2 j respectively, and then perform context information encoding on each of them to obtain two semantic segmentation maps S1 and S2 with the same size as the input image, corresponding to the original images T1 and T2 respectively;

[0038] S4. Combine the semantic segmentation maps S1 and S2 with the change detection map O to obtain the final semantic change detection result maps O1 and O2.

[0039] In this embodiment, the execution network of steps S1 - S4 is simply referred to as OCHRSCD. The execution processes of steps S1 - S4 will be further described in detail below in combination with the structure of OCHRSCD.

[0040] In this embodiment, the siamese high-resolution feature extraction module in step S1 includes two feature extraction branches, and the weights are shared between the two branches.

[0041] In this embodiment, the high-resolution network model used in the feature extraction branch maintains the high-resolution branch, enabling the network to effectively retain the detailed information of the input image.

[0042] In this embodiment, refer to Figure 1 ... The high-resolution network model includes multiple small convolution modules. All the convolution modules in the upper and lower high-resolution network model branches are respectively named convolution module C 1 i,j and convolution module C 2 i,j (i≥1, j≥0). The feature map passes through convolution module C 1 i,j and convolution module C 2 i,j to obtain a new feature map I 1 i,j and I 2 i,j . The resolutions of feature maps I 1 i,j and I 2 i,j are denoted as H 1 i,j ×W 1 i,j and H 2 i,j ×W 2 i,j , and the number of channels is denoted as C 1 i,j and C 2 i,j , where H 1 i,j =H 2 i,j =H input / 2 j , W 1 i,j =W 2 i,j =W input / 2 j , C 1 i,j =C 2 i,j =32×2 j , where H input and W input are the resolution sizes of the input image pair T.

[0043] In this embodiment, optionally, the inputs of convolution modules C 1 i,j and C 1 i,j come from convolution module C1 i-1,y and C 2 i-1,y The output of (where \(i\geq2, y\in[0, i - 2]\)), where C 1 i-1,y The output is denoted as I 1 i-1,y , C 2 i-1,y The output is denoted as I 2 i-1,y ; Separately, the convolutional module C 1 1,0 and C 2 1,0 The input comes from the feature maps obtained by T1 and T2 through convolutional modules C0 and C1: The sizes of T1 and T2 are H input ×W input ×3. By passing through convolutional modules C0 and C1, the number of channels is changed to obtain H input ×W input ×32 feature maps, which are the inputs of the convolutional module C 1 1,0 and C 2 1,0 .

[0044] In this embodiment, in step S1, the convolutional modules C 1 i,j and C 2 i,j There are multiple inputs with different resolutions and numbers of channels, and it is necessary to unify the resolution and the number of channels for addition and fusion. The rules for changing the resolution and the number of channels are as follows: Figure 1 The twin high-resolution feature extraction module in contains three types of arrows. The horizontal arrow represents ordinary convolution, the diagonally upward arrow represents convolution and upsampling operations, and the diagonally downward arrow represents strided convolution. For the input I i-1,y (i≥2, y∈[0, i - 2]), when y < j, perform \(j - y\) strided convolutions with a stride of 2 on the feature map I i-1,y . Each strided convolution doubles the number of channels of the feature map and halves the resolution through a 3×3 convolution, which is represented by a diagonally downward arrow in the Figure 1 feature extraction module; when y = j, perform a 3×3 convolution on the feature map I i-1,y for feature extraction with the number of channels and resolution unchanged, which is represented by a horizontal arrow in the Figure 1 feature extraction module; when y > j, perform a 3×3 convolution on the feature map I i-1,y and change the number of channels to 32×2 j , and use bilinear interpolation upsampling operation to change the resolution to H input / 2 j ×Winput / 2 j , in Figure 1 the feature extraction module, it is represented by a diagonally upward arrow. The convolutional modules C 1 i,j and C 2 i,j After inputting multiple feature maps of and into the same resolution and number of channels, addition is used to fuse the feature maps to obtain I 1 i,j and I 2 i,j .

[0045] In this embodiment, the strided convolution refers to a convolution operation with a stride of 2, which replaces the convolution operation and the pooling operation. Each time the feature map passes through the strided convolution, the number of channels doubles and the resolution is halved.

[0046] In this embodiment, after passing through all the convolutional modules in the feature extraction module in step S1, four feature maps F 1 j ’ and F 2 j ’ are obtained, where j ∈ {0, 1, 2, 3}, and the resolutions and numbers of channels of F 1 j ’ and F 2 j ’ are H input / 2 j ×W input / 2 j and 32×2 j , respectively. Then, F 1 1’, F 1 2’, F 1 3’ are unified in resolution and number of channels to H input ×W input ×32 through 3×3 convolution and bilinear interpolation upsampling operations with factors of two, four, and eight, respectively, to obtain four feature maps with the same resolution and the same number of channels, named F 1 j , where j ∈ {0, 1, 2, 3}; F 2 1’, F 2 2’, F 2 3’ are unified in resolution and number of channels to H input ×W input ×32 through 3×3 convolution and bilinear interpolation upsampling operations with factors of two, four, and eight, respectively, to obtain 4 feature maps with the same resolution and the same number of channels, named F 2 j , where j ∈ {0, 1, 2, 3}.

[0047] In this embodiment, in step S2, the feature pair F 1j and F 2 j Performing differential feature extraction means taking the absolute difference between the corresponding feature pairs F 1 j and F 2 j as the differential feature d i , i ∈ {0, 1, 2, 3}, the differential information feature map d i has a resolution and number of channels both equal to H input ×W input ×32.

[0048] In this embodiment, in step S2, after stacking the channels of d j we obtain a feature map P of size H input ×W input ×128. Then we perform context information encoding on the feature map P to carry out the change detection task between two images. As Figure 2 shown, we first change the number of channels of the feature map P through a simple 1×1 convolution operation to obtain a rough change detection map of size 1×H×W; multiply P by the corresponding rough change detection map matrix to obtain a vector of length 128, which is the object region representation; calculate the relationship matrix R between P and its corresponding 1×128 object region feature representation, with a size of 1×H input ×W input , where the relationship matrix refers to the similarity between pixels and regions; then we perform weighted summation of the object region features according to the values in the relationship matrix R to obtain the context information encoding representation Q, with a size of H input ×W input ×128; stack the original feature map P and the context information encoding representation Q in channels, enhance the information, then connect a convolution layer to change the number of channels of the feature map, and finally obtain the final change detection result map O through post - processing, with a size of H input ×W input ×1.

[0049] In this embodiment, the above - mentioned post - processing refers to performing binaryzation processing using threshold segmentation to obtain the final detection result map O.

[0050] In this embodiment, all the convolution modules included in step S2 are composed of a 3x3 convolution layer, a batch normalization layer, and a rectified linear unit.

[0051] In this embodiment, in step S3, after respectively stacking the channels of the feature pairs F 1 j and F 2 j we obtain a size of H input ×W inputFeature maps P1 and P2 of ×128 are used to encode the context information of feature map P for the semantic segmentation task of the original image. The overall process is similar to the previous step. As Figure 2 shown, first, feature maps P1 and P2 are each passed through a simple 1×1 convolution operation to change their number of channels, respectively obtaining rough semantic segmentation maps of size 1×H×W; multiplying P1 and P2 with the corresponding rough semantic segmentation map matrices respectively yields 1 vector of length 128, that is, the object region representation; calculating the relationship matrices R1 and R2 between P1 and P2 and their corresponding 1×128 object region feature representations, with a size of 1×H input ×W input ; then, according to the values of R1 and R2 in the relationship matrix, the object region features are weighted and summed to obtain the context information encoded representations Q1 and Q2 of P1 and P2 respectively, with a size of H input ×W input ×128; after stacking the channel information of the original feature maps P1 and P2 with their corresponding context information encoded representations Q1 and Q2 and enhancing it, a convolution layer is connected to change the number of channels of the feature map, respectively obtaining the semantic segmentation maps S1 and S2 of the original image pair, with a size of H input ×W input ×1.

[0052] In this embodiment, combining the semantic segmentation maps S1 and S2 with the change detection map O in step S4 means only retaining the semantic segmentation results of the regions in the semantic segmentation maps S1 and S2 that are the same as the changed regions in the change detection map O, and ignoring the semantic segmentation results of the regions in O that have not changed, to obtain the final semantic change detection result maps O1 and O2.

[0053] In this embodiment, after step S4, there is also a step of using the deep supervision method to guide the prediction network. When training the network using the deep supervision method, the binary cross-entropy loss L is obtained by using the change binary map O and the true change label bce , and the cross-entropy losses are obtained by using the semantic change detection result maps O1 and O2 and the true semantic change labels respectively and respectively. After that and L bce are weighted and added to obtain the total loss L. L is backpropagated and iterated repeatedly until the number of iterations reaches the set initial value, at which point the training is determined to be completed.

[0054] L bce is as follows:

[0055]

[0056] where n represents the number of pixels in the image, y i represents the true change map of the building, y i∈{0,1} represents the value at position i in y, where 1 indicates that this pixel has changed, and 0 indicates that this pixel has not changed. x i represents the predicted change map output by the network model, x i ∈[0,1] represents the value at position i in x, representing the probability that the predicted pixel point has changed.

[0057] and represents calculating the cross-entropy loss L between the semantic change detection result maps O1 and O2 and their corresponding semantic change detection labels respectively ce . L ce is as follows:

[0058]

[0059] where Class represents the number of categories of semantic segmentation, p i represents the true label at position i, q i represents the predicted value at position i.

[0060] In this embodiment, and L bce After weighted addition, the total loss L is obtained, where the weights of and are 1, and the weight of L bce is 2. L is as follows:

[0061]

[0062] To verify the effectiveness of OCHRSCD, in this embodiment, the public dataset SECOND is used for training and testing the network framework, and comparisons are made with other methods. The SECOND dataset contains 2968 groups of training data and 647 groups of test data. Each group of data contains two images taken at different times, and the size of each image is 512×512 pixels.

[0063] The algorithm proposed in this embodiment is compared with five latest semantic change detection methods: DSCD (Direct Segmentation Change Detection), SCDS (Separate Change Detection and Segmentation), ICDS (Intergrated Change Detection and Segmentation), HBSCD (HRNet based semantic change detection), and SCDNet (Segmentation Change Detection Network). The specific results are as Figure 3As shown. There are three evaluation metrics, namely mean Intersection over Union (mIoU), Separate Kappa (SeK), and comprehensive score Score. Combining Figure 3 It can be seen that the three evaluation metrics of the method OCHRSCD in this embodiment are all the best results and reach the highest Score (39.31%). Compared with the second-best (SCDNet), OCHRSCD improves the accuracy of Score by 0.83%, mIoU by 0.15%, and SeK by 1.12%. Figure 4 This is a schematic diagram of the semantic change detection results of the method in this embodiment on three pairs of data. The second row represents the label maps corresponding to the T1 image and the T2 image, and the third row represents the results obtained by OCHRSCD. The figure shows that compared with the label maps, the detection effect of the changed regions is better, the contours of the changed regions in the detection results are clear, and there is no adhesion in the detection results of the regions with dense changed distributions. And the accuracy of semantic segmentation for the changed regions is also good.

[0064] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A semantic change detection method based on a high-resolution convolutional network and context information encoding, characterized in that The steps are as follows: S1. Respectively input remote sensing images T1 and T2 at different time phases into the twin high-resolution feature extraction module after changing the number of channels through convolution modules C0 and C1, and perform feature extraction and exchange of feature information at different scales through multiple small convolution modules in the feature extraction module to obtain feature maps F 1 j ’ and F 2 j ’; Unify the resolution and number of channels of the features F 1 j ’ and F 2 j ’ respectively through convolution and upsampling operations to obtain new feature information F 1 j and F 2 j ; S2, pair the feature F 1 j and F 2 j to perform differential feature extraction to obtain the differential information feature d j between the two images. Stack the d j channels and perform context information encoding. After post-processing, a binary change map O with the same size as the input image is obtained; S3. Stack the feature pairs F 1 j and F 2 j respectively in channels, then perform context information encoding on each of them to obtain two semantic segmentation maps S1 and S2 with the same size as the input image, corresponding to the original images T1 and T2 respectively; S4. Combine the semantic segmentation maps S1 and S2 with the change detection map O to obtain the final semantic change detection result maps O1 and O2.

2. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, wherein: In step S1, the Siamese high-resolution feature extraction module includes two feature extraction branches, and the weights are shared between the two branches.

3. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 2, wherein: The high-resolution network model includes multiple small convolutional modules, and all convolutional modules in the upper and lower high-resolution network model branches are respectively named convolutional module C 1 i,j and convolutional module C 2 i,j , where i≥1, j≥0, and the feature map passes through convolutional module C 1 i,j and convolutional module C 2 i,j to obtain a new feature map I 1 i,j and I 2 i,j , and the resolutions of feature map I 1 i,j and I 2 i,j are denoted as H 1 i,j ×W 1 i,j and H 2 i,j ×W 2 i,j , and the number of channels is denoted as C 1 i,j and C 2 i,j , where H 1 i,j =H 2 i,j =H input / 2 j , W 1 i,j =W 2 i,j =W input / 2 j , C 1 i,j =C 2 i,j =32×2 j , where H input and W input are the resolution sizes of the input image pair T 4. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 3, wherein: Convolution module C in step S1 1 i,j and C 1 i,j take the input from the outputs of convolution modules C 1 i-1,y and C 2 i-1,y where i ≥ 2, y ∈ [0, i - 2], and the output of C 1 i-1,y is denoted as I 1 i-1,y , and the output of C 2 i-1,y is denoted as I 2 i-1,y ; separately, the inputs of convolution modules C 1 1,0 and C 2 1,0 come from the feature maps obtained by T1 and T2 through convolution module C0 and convolution module C1.

5. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, wherein: Convolution module C in step S1 1 i,j and C 2 i,j There are multiple inputs with different resolutions and numbers of channels, and the resolutions and numbers of channels need to be unified for addition and fusion; the rules for changing the resolutions and numbers of channels are as follows: for input I i-1,y (i≥2,y∈[0,i - 2]), when y < j, perform (j - y) strided convolutions with a stride of 2 on the feature map I i-1,y Each strided convolution doubles the number of channels of the feature map and halves the resolution through a 3×3 convolution; When y = j, perform a 3×3 convolution on the feature map I i-1,y to extract features while keeping the number of channels and resolution unchanged; when y > j, perform a 3×3 convolution on the feature map I i-1,y and change the number of channels to 32×2 j , and use bilinear interpolation upsampling operation to change the resolution to H input / 2 j ×W input / 2 j; Input multiple feature maps of the convolution modules C 1 i,j and C 2 i,j with unified resolution and number of channels, and then fuse the feature maps using addition.

6. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, wherein: In step S2, for the feature pair F 1 j and F 2 j performing differential feature extraction means taking the absolute difference between the corresponding feature pairs F 1 j and F 2 j as the differential feature d i .

7. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, characterized in that: In step S2, the post-processing refers to performing binaryzation processing using threshold segmentation to obtain the final detection result map O.

8. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, characterized in that: In step S4, combining the semantic segmentation maps S1 and S2 with the change detection map O means only retaining the semantic segmentation results of the regions in the semantic segmentation maps S1 and S2 that are the same as the changed regions in the change detection map O, and ignoring the semantic segmentation results of the regions in O that have not changed.

9. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, wherein: After step S4, there is also a step of guiding the prediction network using the deep supervision method. When training the network using the deep supervision method, the binary change map O and the true change label are used to obtain the binary cross-entropy loss Lbce, and the semantic change detection result maps O1 and O2 and the true semantic change labels are used to obtain the cross-entropy losses L1ce and L2ce respectively; L1ce, L2ce, and Lbce are weighted and added together to obtain the total loss L, and L is backpropagated and iterated repeatedly until the number of iterations reaches the set initial value, at which point it is determined that the training is completed.

10. The semantic change detection method based on a high-resolution convolutional network and context information encoding according to claim 1, wherein: All convolutional modules are composed of a 3×3 convolutional layer, a batch normalization layer, and a rectified linear unit.

Citation Information

Patent Citations

  • Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field

    AU2020103901A4

  • Remote sensing image road network extraction method based on multi-scale feature fusion

    CN113850824A