A multi-branch collaborative semantic change detection method, system, device and medium

Through the multi-branch collaborative semantic change detection method, pseudo-label, dual-resolution network and context information interaction modules are used to solve the problem of data set distribution characteristics and categories imbalance in remote sensing images, and more efficient semantic change detection is achieved.

CN116758366BActive Publication Date: 2025-07-11TIANJIN JIAOJIANYAN INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310538911.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-12
Publication Date
2025-07-11
Estimated Expiration
2043-05-12

AI Technical Summary

Technical Problem

The existing semantic change detection methods have problems in remote sensing images with complex distribution characteristics of data sets, spectral variability, noise interference, category imbalance and insufficient interaction between subtasks, resulting in poor detection results, low feature utilization rate, and serious error accumulation.

Method used

Multi-branch collaborative semantic change detection method is adopted, by introducing pseudo-labels to increase the number of labeled pixels, adding high-resolution branches for dual-resolution networks, introducing a context information interaction module to enhance context information representation, and feature fusion is performed through the channel attention mechanism to improve the detection accuracy of changing areas.

Benefits of technology

It effectively alleviates the difficult sample separation problem in complex backgrounds, reduces error accumulation, improves detection accuracy and feature utilization, and improves the performance of semantic change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758366B_ABST
    Figure CN116758366B_ABST
Patent Text Reader

Abstract

A multi-branch collaborative semantic change detection method, system, device and medium. The method is as follows: introducing pseudo-labels to predict semantic class labels for unchanged regions; implementing a dual-resolution network by adding an additional high-resolution branch to ResNet; introducing a context information interaction module to enhance the context information representation of the target region using the target region representation, and performing a splicing operation on the original feature and the context information interaction feature to obtain the final enhanced feature; performing channel fusion on the extracted enhanced feature in multiple ways, and introducing a channel attention mechanism before dimensionality reduction to focus on important features in the channel domain, so as to balance complexity and feature richness, improve the detection accuracy of changed regions, and reduce the false detection rate of unchanged regions; the system, device and medium implement multi-branch collaborative semantic change detection based on the above method; the present invention effectively models ground object coverage information and change information using a multi-branch network structure, and through multi-branch interaction and collaboration, improves the feature utilization rate and jointly enhances the task performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of semantic change detection, and in particular relates to a multi-branch collaborative semantic change detection method, system, device and medium. Background Art

[0002] When detecting a given pair of remotely sensed images of the same area at different times, most existing algorithms focus on detecting the positions of changed and unchanged pixel points between the images at different times, that is, binary change detection (BCD), which provides relatively single information and cannot provide richer and more refined semantic information for subsequent applications; semantic change detection (SCD) can not only identify the changed areas, but also obtain the change categories of the bi-temporal areas. The semantic change detection task can be decomposed into the combination of two tasks: semantic segmentation of the pair of remotely sensed images and binary change detection.

[0003] Due to the limitation of large-scale semantic change detection datasets, most existing semantic change detection methods focus on scene-level changes, where the semantic change map is only generated by rough boundaries or scarce category information. With the continuous emergence of pixel-level semantic change detection datasets, some specific models and methods for semantic change detection have also been proposed. However, in the semantic change detection task, there are 1) complex background: the regions of the bi-temporal remotely sensed image pair have complex surface feature distributions, specifically manifested as inconsistent resolutions, no obvious regular distribution between different images, and many difficult-to-separate samples, etc., which have a negative impact on the semantic change detection results; 2) existence of "salt-and-pepper noise": the bi-temporal remotely sensed image pair has spectral variability, and the regions identified by traditional pixel classification methods are often not fine and present identification results similar to salt-and-pepper noise; 3) visual features are easily confused: there are differences in imaging conditions between the bi-temporal remotely sensed image pair, or there is noise interference, and the features in the image pair may be visually confused compared with the original situation, increasing the difficulty of identifying the changed areas; 4) class imbalance: on the basis of traditional change detection to identify changed and unchanged regions, semantic change detection also needs to provide detailed land cover categories before and after observation. There are class imbalance problems between the changed class and the unchanged class, and among the changed classes, it is difficult to accurately mine the classes with few training samples; 5) insufficient interaction between sub-tasks: there is easy error accumulation between the sub-tasks decomposed from the semantic change detection task, the interaction of the sub-task branches is limited, and the feature utilization rate is insufficient, resulting in poor detection effects; and other challenges, resulting in error accumulation and low feature utilization rate, making the prediction model unable to learn sufficient and effective features, thereby affecting the detection performance of the model.

[0004] Existing semantic change detection method solutions include algorithms based on traditional manual design and algorithms based on deep learning.

[0005] In the algorithms based on traditional manual design, domestic and foreign scholars mainly conduct manual design on change detection algorithms. These methods require prior knowledge, such as texture, morphology, etc. Representative algorithms include support vector machine (SVM), decision tree (DT), random forest, Markov random field (MRF), etc. [Sui Haigang, Feng Wenqing, Li Wenzhuo, et al. A review of multi-temporal remote sensing image change detection methods [J]. Journal of Wuhan University, Information Science Edition, 2018, 43(12): 1885-1898.]. However, it is difficult for manually designed algorithms to capture high-level semantic information, with slow speed, complex design process, and poor generalization. In recent years, with the rapid development of artificial intelligence and image processing technologies, deep learning methods have been quickly applied to the field of change detection. Among them, due to the ability to extract both low-level detail information and high-level semantic features simultaneously, Convolutional Neural Networks (CNN) have been proven to have good performance in change detection tasks [Lu D, Cheng S, Wang L, et al. Multi-scale feature progressive fusion network for remote sensing image change detection [J]. Scientific Reports, 2022, 12(1): 11968.], [Shi W, Zhang M, Zhang R, et al. Change detection based on artificial intelligence: State-of-the-art and challenges [J]. Remote Sensing, 2020, 12(10): 1688.].

[0006] However, the semantic information obtained from the change detection task is not rich enough, and there are significant limitations in many application scenarios. The research [Bovolo F, Bruzzone L. The time variable in data fusion: A change detection perspective[J]. IEEE Geoscience and Remote Sensing Magazine, 2015, 3(3): 8-26.] attempts to introduce scene-level semantic labels into the change detection task. The research [Wu C, Zhang L, Zhang L. A scene change detection framework for multi-temporal very high resolution remote sensing images[J]. Signal Processing, 2016, 124: 184-197.] proposes a strategy of classifying first and then comparing, and proposes a scene-level CD framework, but it still extracts features manually. The research [Ru L, Wu C, Du B, et al. Deep canonical correlation analysis network for scene change detection of multi-temporal VHR imagery[C] / / 2019 10th International Workshop on the Analysis of Multitemporal Remote Sensing Images (MultiTemp). IEEE, 2019: 1-4.] introduces CNN and proposes a deep learning-based scene-level change detection framework. However, scene-level change detection represents target objects with rectangular regions, lacking boundary information of the objects. The algorithm has low accuracy and also limits its application in real scenarios. Therefore, it is urgent to solve the problem of semantic change detection based on pixel-level information.

[0007] The pixel-level semantic change detection task requires pixel-level semantic annotation of the input image [Peng D, Bruzzone L, Zhang Y, et al. SCDNET: A novel convolutional network for semantic change detection in high resolution optical remote sensing imagery [J]. International Journal of Applied Earth Observation and Geoinformation, 2021, 103: 102465.]. To utilize the correlation of the dual-temporal image pair, most of the existing semantic change detection methods are based on the architecture of the siamese network, that is, the network structures are the same, the weights are shared, and good results have been achieved. In the recent research on semantic change detection, the research [Daudt R C, Le Saux B, Boulch A. Fully convolutional siamese networks for change detection [C] / / 2018 25th IEEE International Conference on Image Processing (ICIP). IEEE, 2018: 4063-4067.] proposed three fully convolutional network architectures and designed a segmentation branch and a change branch. In the research [Yang K, Xia G S, Liu Z, et al. Asymmetric siamese networks for semantic change detection in aerial images [J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 1-18.], a three-branch CNN introduced a gating unit in the decoder to improve the feature representation. However, in the above research, the sub-tasks are modeled separately without considering the correlation, which is prone to error accumulation.The research [Zhu Q, Guo X, Deng W, et al. Land-use / land-cover change detection based on a Siamese global learning framework for high spatial resolution remote sensing imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 184: 63-78.] proposed a CNN architecture to complete the semantic change detection task. The encoder uses a Siamese network to extract features, and a decoder is designed to detect change information. The research [Wang D, Zhao F, Wang C, et al. Y-Net: A multiclass change detection network for bi-temporal remote sensing images[J]. International Journal of Remote Sensing, 2022, 43(2): 565-592.] introduced an attention design in the decoder to enhance features. The research [Ding L, Guo H, Liu S, et al. Bi-temporal semantic reasoning for the semantic change detection in HR remote sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 1-14.] introduced a fusion unit to model the temporal correlation between bi-temporal image pairs. However, there are still the following defects:

[0008] 1) The distribution characteristics of the dataset are not considered, and there are problems with difficult-to-separate samples in complex backgrounds, making it difficult to mine effective high-level semantic information and obtain representative deep features;

[0009] 2) The spectral variability of remote sensing images is not considered, and the relationship between pixels and space is not considered in the context of rich and complex geometric information, that is, the interaction of context information is not considered;

[0010] 3) The problem of sample imbalance is not considered, resulting in greater segmentation difficulty for some classes. The network will fit towards the samples with more classes, resulting in inaccurate semantic change detection results;

[0011] 4) The correlation of hidden deep features is not mined, resulting in insufficient feature utilization and poor detection effect. Summary of the Invention

[0012] In order to overcome the above-mentioned deficiencies of the prior art, the purpose of the present invention is to propose a multi-branch collaborative semantic change detection method, system, device and medium. Pseudo-labels are introduced to predict semantic class labels for unchanged regions, thereby increasing the number of labeled pixels and alleviating the problem of small data volume. A dual-resolution network is implemented by adding an additional high-resolution branch to ResNet. Then, a context information interaction module is introduced to enhance its context information representation using the target region representation, explicitly transforming the pixel classification problem into an object region classification problem. The extracted enhanced features are then subjected to channel fusion in various ways. At the same time, in order to balance complexity and feature richness, a channel attention mechanism is introduced before dimensionality reduction to focus on important features in the channel domain, thereby improving the detection accuracy of changed regions and reducing the false detection rate of unchanged regions. The present invention effectively models ground object coverage information and change information using a multi-branch network structure, provides more consistent change information for sub-tasks by embedding more object-oriented features, and the multi-branches interact and cooperate to improve feature utilization and jointly enhance the performance of the detection task.

[0013] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0014] A multi-branch collaborative semantic change detection system and method, which introduce pseudo-labels to predict semantic class labels for unchanged regions, thereby increasing the number of labeled pixels; implement a dual-resolution network by adding an additional high-resolution branch to ResNet; then, introduce a context information interaction module to enhance its context information representation using the target region representation to obtain context information interaction features, and perform a splicing operation on the original features and the context information interaction features to obtain the final enhanced features; then, perform channel fusion on the extracted enhanced features in various ways, and at the same time introduce a channel attention mechanism before dimensionality reduction to focus on important features in the channel domain, balancing complexity and feature richness to improve the detection accuracy of changed regions and reduce the false detection rate of unchanged regions.

[0015] A multi-branch collaborative semantic change detection method specifically includes the following steps:

[0016] S1. Divide the semantic change detection sample set, and randomly divide the publicly available semantic change detection data set into a training set and a test set;

[0017] S2. Read the images in the training set divided in step S1, and perform data processing and data augmentation on them:

[0018] S3. Use a multi-resolution convolutional neural bidirectional fusion network to extract features from the dual-temporal remote sensing images. The extracted multi-resolution features are X ms ,

[0019] Add an additional high-resolution branch on the basis of the classification backbone network ResNet. The ResNet network is divided into four stages when performing downsampling. The output result of each stage halves the size of the feature map and doubles the number of channels, gradually extracting high-level semantic features. At different stages, "high-low fusion" and "low-high fusion" are respectively performed on the features of the high-resolution branch and the features of the low-resolution branch. Through bidirectional feature fusion, make full use of the multi-resolution spatial information and semantic information;

[0020] S4. Introduce the idea of class self-attention mechanism to construct a context information interaction module, and use the context information interaction module to enhance the context information. Finally, the enhanced features after context information interaction are X ce ,

[0021] S5. Perform channel fusion on the features extracted from the dual-temporal phase. Finally, the fused features are X fus ,

[0022] S6. According to the finally fused features in steps S4 and S5, through the corresponding semantic segmentation decoder and binary change detection decoder, obtain the final output result.

[0023] The specific method of step S2 is as follows:

[0024] First, select multiple classic network models, including ResNet101 and HRNet-W64. Input the original data into the model for training. Use the trained model to predict the specific semantic categories of the unchanged regions in the semantic change detection dataset through the model integration method. Take the predicted results of the unchanged regions as pseudo-labels, and combine them with the existing labels of the changed regions to achieve secondary annotation of the dataset and increase the number of labeled pixels;

[0025] Then, perform data augmentation on the data by randomly flipping left and right, flipping up and down, and rotating at multiple angles.

[0026] The model used for pseudo-label generation in step S2 does not participate in subsequent training and testing.

[0027] The specific method of step S4 is as follows:

[0028] S401. Perform convolution and dimension conversion operations on the multi-resolution features X extracted in step S3 ms See formula (1):

[0029] X feats = RELU(BN(Conv 3×3 (X ms ))) (1)

[0030] X permute = Permute(X feats ) (2)

[0031] where Permute(·) refers to the dimension permutation operation;

[0032] S402. Obtain the rough semantic segmentation result RSS for X ms through a simple rough segmentation network;

[0033] RSS = Conv 1×1 (RELU(BN(Conv 1×1 (X ms )))) (3)

[0034] S403. Perform Softmax and dimension conversion operations on the rough semantic segmentation result RSS obtained in step S402 to obtain C groups, where C is the segmentation category, and the vector set V = {V i , i = 1,..., C}, and each vector V i is the feature representation of a category;

[0035] X soft = sofmtax(Permute(X ms )) (4)

[0036] S404. Perform matrix multiplication on the results obtained in steps S401 and S403 to obtain the object context representation;

[0037] X object = Matmul(X permute , X soft ) (5)

[0038] S405. Calculate the relationship between each pixel and the object context representation obtained in step S404 to obtain the final representation of each pixel;

[0039] Introduce Query, Key, Value, and use the idea of the self-attention mechanism to obtain the final context information representation:

[0040] Query = Permute(RELU(BN(Conv 1×1 (RELU(BN(Conv 1×1 (X feats )))))) (6)

[0041] Key = RELU(BN(Conv 1×1 (RELU(BN(Conv 1×1 (X object ))))) (7)

[0042] Value = Permute(RELU(BN(Conv 1×1 (X object )))) (8)

[0043]

[0044] X c That is, the context information interaction feature. Concatenate it with the original feature to obtain the final enhanced feature:

[0045] X ce = CON(X c , X feats ) (10)

[0046] where L is the number of channels of the last convolutional layer in this module;

[0047] The enhanced feature after the final context information interaction is X ce ,

[0048] The specific method of step S5 is as follows:

[0049] S501. Perform channel fusion on the enhanced feature extracted in step S4 in multiple ways, including feature addition, subtraction, and concatenation. Provide the features obtained by various combinations to the binary change detection network branch. See formulas (11)-(13):

[0050]

[0051]

[0052]

[0053] S502. Introduce a channel attention mechanism to process the feature channels. The features obtained by adding, subtracting, and concatenating along the channel dimension in formulas (11)-(13) are operated as follows:

[0054] 1) Squeeze operation. Through global average pooling, perform feature compression along the spatial dimension, turning each two-dimensional feature channel into a real number. This real number has a global receptive field to some extent, and the output dimension matches the number of input feature channels;

[0055] 2) Incentive operation, which generates weights for each feature channel through parameters;

[0056] 3) Reweighting operation. Consider the weights output by the incentive operation as the importance of each feature channel after feature selection, and then multiply and weight each channel to all the features before step 3) of step S502 one by one, completing the recalibration of the original features in the channel dimension. The introduced channel attention mechanism selects more important features through the above operations, and then performs dimensionality reduction operation to compromise between the richness of features and the computational complexity;

[0057] X fus = SA(X con ) (14)

[0058] where SA(·) refers to the channel attention operation;

[0059] The finally fused feature is X fus ,

[0060] The specific method of step S6 is as follows:

[0061] The features extracted in step S4 and step S5 are passed through the corresponding semantic segmentation decoder and the binary change detection decoder to obtain the final output results: a binary change map CM, and two semantic change detection maps SCD1 and SCD2;

[0062]

[0063]

[0064]

[0065] where Mask(·) is a dot product operation. After multiplying the semantic segmentation result and the binary change detection map CM, only the semantic information of the changed area is retained; The structures of the semantic segmentation decoder and the binary change detection decoder are simplified fully convolutional networks, which obtain segmentation and change results through convolution, BN, and Relu, and then perform interpolation to obtain a semantic segmentation map and a binary change map with the same size as the input.

[0066] A multi-branch collaborative semantic change detection system includes

[0067] A multi-resolution convolutional neural network encoder, which is a twin structure with the same structure and shared weights, and is used to input dual-temporal images I1, I2 and extract multi-resolution depth features;

[0068] The context information interaction module is a twin structure with two identical structures sharing weights, which is used to perform context information interaction on the target area, splice it with the original features, and obtain the final enhanced features;

[0069] The channel fusion module is used to explore the correlation of hidden deep features, enhance the collaborative interaction between subtasks, provide richer and more effective features for the binary change detection network branch, and improve the feature utilization rate;

[0070] The semantic segmentation decoder is a twin structure with two identical structures sharing weights, which is used to output two semantic segmentation maps and

[0071] The binary change detection decoder is used to output a binary change map CM, as well as two semantic change detection maps SCD1 and SCD2.

[0072] A detection device based on a multi-branch collaborative semantic change detection method, comprising:

[0073] A memory for storing computer programs, data and models;

[0074] A processor for implementing the multi-branch collaborative semantic change detection method according to any one of steps S1 to S6 when executing the computer program.

[0075] A computer-readable storage medium is responsible for reading and storing programs and data. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can perform multi-branch collaborative semantic detection based on the detection method according to any one of steps S1 to S6.

[0076] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0077] 1. The deep features extracted by the multi-resolution fusion strategy proposed by the present invention are more effective.

[0078] The prior art does not consider the distribution characteristics of the dataset and the problem of difficult-to-separate samples in complex backgrounds, that is, it ignores the necessity of multi-resolution. In the present invention, a multi-resolution fusion strategy is introduced in the encoder to extract deep features, effectively fuse feature maps of different resolutions on the basis of expanding the effective receptive field, covering both low-level detail information and high-level semantic information, capturing subtle ground changes in the scene, obtaining richer and more effective deep features, reflecting the trade-off between accuracy and speed, alleviating the problem of difficult-to-separate samples in complex backgrounds, further reducing the error accumulation between subtasks, and finally improving the detection ability of the network for different types and sizes of targets.

[0079] 2. The context information interaction proposed by the present invention is more effective.

[0080] The prior art does not take into account the spectral variability of remote sensing images. The regions identified in rich and complex geometric backgrounds are often not fine and contain noise. Therefore, the context information interaction of the dual-temporal image itself is necessary. A pixel itself does not have semantics, and its semantics are determined by the overall image or the target region. Therefore, it is highly dependent on the context. The context information interaction module proposed by the present invention can effectively alleviate the "salt and pepper" noise problem of the semantic change detection results, and to a certain extent, also alleviates the problem of difficult-to-separate samples in complex backgrounds, achieving higher detection accuracy.

[0081] 3. The channel fusion method proposed by the present invention is more effective.

[0082] The prior art does not consider the deep interaction between subtasks and the correlation of hidden deep features, resulting in insufficient feature utilization. Moreover, the change detection network of the prior art does not consider the task itself during design. The present invention uses multiple methods to fuse channel features to obtain more change features, and then uses channel attention to screen and fuse the features, highlighting valuable deep features, enhancing the collaborative interaction between subtasks, suppressing redundant information, improving the detection effect, increasing the utilization rate of features, and at the same time, the proposed attention fusion module accelerates the model convergence during the training process.

[0083] 4. The combination method of pseudo-label and weighted cross-entropy loss function proposed by the present invention is more effective.

[0084] The prior art does not consider problems such as less sample data volume, imbalance, and difficult-to-separate samples. The semantic change detection dataset used in the present invention only has accurate labels for the changed regions, and there are problems of extremely unbalanced samples and difficult-to-separate samples, resulting in the network fitting towards the samples with more categories, increasing the difficulty of semantic change detection and the result not being accurate enough. The present invention performs pseudo-label annotation on the unchanged regions, gives greater weights to the samples with fewer categories and difficult to separate, and uses the combination of pseudo-labels and weighted cross-entropy loss to alleviate the problems of less data volume, imbalance, and difficult-to-separate samples to a certain extent.

[0085] 5. The model in the present invention adopts a modular design concept. The basic network structures of feature extraction and feature fusion can both be replaced. Modules can be added or modified according to the deficiencies of the basic network in the semantic change detection task. With the development of emerging technologies and the proposal of better network modules, the present invention can be iteratively updated at any time to improve the model performance.

[0086] In summary, compared with the prior art, the present invention achieves better detection performance and realizes a good compromise between accuracy and model complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 This is the overall flowchart of the present invention.

[0088] Figure 2 This is the flowchart of the context information interaction module in the present invention.

[0089] Figure 3 This is the flowchart of the channel fusion unit based on the attention mechanism in the present invention.

[0090] Figure 4 This is a set of experimental result diagrams of the present invention.

[0091] Figure 5 is Figure 4 the corresponding true label map. Detailed implementation manners

[0092] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0093] A multi-branch collaborative semantic change detection method, see Figure 1 , which specifically includes the following steps:

[0094] S1. Divide the semantic change detection sample set, and randomly divide the publicly available semantic change detection data set into a training set and a test set;

[0095] S2. Read the images in the training set divided in step S1, and perform data processing and data augmentation on them:

[0096] The corresponding regions of the labels in the semantic change detection data set can be divided into changed regions and unchanged regions. The semantic change detection data set used in this experiment only provides the labels of the changed regions. Therefore, to a certain extent, it affects the accuracy of the sub-task during feature extraction. The present invention solves this problem in the form of pseudo-labels: Since models with high complexity and large number of parameters generally have better performance and stronger generalization ability, several classic network models with high complexity and large number of parameters are selected, such as ResNet101, HRNet-W64, etc. The original data is input into the model for training, and the trained model is used to predict the specific semantic categories of the unchanged regions in the semantic change detection data set through model integration. The predicted results of the unchanged regions are used as pseudo-labels, combined with the existing labels of the changed regions, to achieve secondary annotation of the data set and increase the number of labeled pixels; it alleviates the problem of small data volume, improves the accuracy of the sub-task at the same time, and the models with high complexity and large number of parameters used for pseudo-label generation do not participate in subsequent training and testing, improving the overall efficiency of the invention;

[0097] Data augmentation is performed on the data by randomly flipping left and right, flipping up and down, rotating 90 degrees, rotating 180 degrees, and rotating 270 degrees.

[0098] S3. Use a multi-resolution convolutional neural bidirectional fusion network to extract features from the dual-temporal remote sensing image pair. The extracted multi-resolution features are X ms ,

[0099] An additional high-resolution branch is added on the basis of the classification backbone network ResNet. The ResNet network can be divided into four stages when performing downsampling. The output result of each stage halves the size of the feature map and doubles the number of channels, gradually extracting high-level semantic features. At different stages, "high-low fusion" and "low-high fusion" are respectively performed on the features of the high-resolution branch and the features of the low-resolution branch. Through bidirectional feature fusion, the spatial information and semantic information of multiple resolutions are fully utilized.

[0100] S4. Introduce the idea of class self-attention mechanism to construct a context information interaction module. See Figure 2 . Use the context information interaction module to enhance the context information. Finally, the enhanced features after context information interaction are X ce ,

[0101] S401. Perform convolution and dimension conversion operations on the multi-resolution features X ms extracted in step S3, as shown in formula (1)

[0102] X feats = RELU(BN(Conv 3×3 (X ms ))) (1)

[0103] X permute = Permute(X feats ) (2)

[0104] where Permute(·) refers to the dimension swapping operation.

[0105] S402. Obtain the rough semantic segmentation result RSS for X ms through a simple rough segmentation network.

[0106] RSS = Conv 1×1 (RELU(BN(Conv 1×1 (X ms )))) (3)

[0107] S403. Perform Softmax and dimensionality conversion operations on the rough semantic segmentation result RSS obtained in step S402 to obtain C groups, where C is the segmentation category, and the vector set V = {V i , i = 1,..., C}, and each vector V i is a feature representation of a category;

[0108] X soft = softmax(Permute(X ms )) (4)

[0109] S404. Multiply the results obtained in steps S401 and S403 in matrix form to obtain the object context representation;

[0110] X object = Matmul(X permute , X soft ) (5)

[0111] S405. Calculate the relationship between each pixel and the object context representation obtained in step S404 to obtain the final representation of each pixel;

[0112] Introduce Quert, Key, and Value, and use the idea of the self-attention mechanism to obtain the final context information representation:

[0113] Query = Permute(RELU(BN(Conv 1×1 (RELU(BN(Conv 1×1 (X feats )))))) (6)

[0114] Key = RELU(BN(Conv 1×1 (RELU(BN(Conv 1×1 (X object ))))) (7)

[0115] Value = Permute(RELU(BN(Conv 1×1 (X object )))) (8)

[0116]

[0117] X c is the context information interaction feature. Concatenate it with the original feature to obtain the final enhanced feature:

[0118] X ce = CON(X c , X feats ) (10)

[0119] where L is the number of channels of the last convolutional layer in this module;

[0120] The enhanced feature after the final context information interaction is X ce ,

[0121] S5. Channel fusion is performed on the features extracted in the two temporal phases. See Figure 3 , and the finally fused feature is X fus ,

[0122] S501. Channel fusion in multiple ways is performed on the enhanced features extracted in step S4 , including three ways: feature addition, subtraction, and concatenation, providing features obtained by various combinations for the binary change detection network branch. See formulas (11)-(13):

[0123]

[0124]

[0125]

[0126] S502. A channel attention mechanism is introduced to process the feature channels. Along the channel dimension, different features are obtained through three ways: feature addition, subtraction, and concatenation in formulas (11)-(13). The following operations are performed on the obtained features:

[0127] 1) Squeeze operation: Through global average pooling, feature compression is performed along the spatial dimension, turning each two-dimensional feature channel into a real number. This real number has a global receptive field to some extent, and the output dimension matches the number of input feature channels;

[0128] 2) Excitation operation: Weights are generated for each feature channel through parameters;

[0129] 3) Reweighting operation: The weights output by the excitation operation are regarded as the importance of each feature channel after feature selection, and then multiplied channel by channel and weighted to all the features before step 3) of S502, completing the recalibration of the original features in the channel dimension. The introduced channel attention mechanism selects more important features through the above operations, and then performs a dimensionality reduction operation to balance the richness of the features and the computational complexity;

[0130] X fus = SA(X con ) (14)

[0131] where SA(·) refers to the channel attention operation;

[0132] The finally fused feature is X fus ,

[0133] S6. According to the finally fused feature in steps S4 and S5, through the corresponding semantic segmentation decoder and binary change detection decoder, obtain the final output result:

[0134] Pass the features extracted in steps S4 and S5 through the corresponding semantic segmentation decoder and the binary change detection decoder to obtain the final output result: a binary change map CM, and two semantic change detection maps SCD1 and SCD2;

[0135]

[0136]

[0137]

[0138] Among them, Mask(·) is a dot product operation. After dot multiplying the semantic segmentation result and the binary change detection map CM, only retain the semantic information of the changed area; the semantic segmentation decoder and the binary change detection decoder have the structure of a simplified fully convolutional network. Obtain the segmentation and change results through convolution, BN, and Relu, and then perform interpolation to obtain a semantic segmentation map and a binary change map with the same size as the input.

[0139] A multi-branch collaborative semantic change detection system includes

[0140] A multi-resolution convolutional neural network encoder, which is a twin structure with the same structure and shared weights, and is used to input dual-temporal images I1, I2 and extract multi-resolution depth features;

[0141] A context information interaction module, which is a twin structure with shared weights and the same structure, and is used to perform context information interaction on the target area and splice it with the original features to obtain the final enhanced features;

[0142] A channel fusion module, which is used to mine the correlation of hidden deep features, enhance the collaborative interaction between subtasks, provide richer and more effective features for the binary change detection network branch, and improve the feature utilization rate;

[0143] A semantic segmentation decoder, which is a twin structure with shared weights and the same structure, and is used to output two semantic segmentation maps and

[0144] A binary change detection decoder for outputting a binary change map CM and two semantic change detection maps SCD1 and SCD2.

[0145] A detection device based on a multi-branch collaborative semantic change detection method, comprising:

[0146] A memory for storing computer programs, data and models;

[0147] A processor for implementing the multi-branch collaborative semantic change detection method according to any one of steps S1 to S6 when executing the computer program.

[0148] A computer-readable storage medium for reading and storing programs and data, the computer-readable storage medium storing a computer program, and the computer program being capable of performing multi-branch collaborative semantic detection based on the detection method according to any one of steps S1 to S6 when executed by a processor.

[0149] For the experimental results of the present invention, see Figure 4 and for the corresponding label map, see Figure 5 . Figure 4 In, I1 and I2 are the input dual-temporal remote sensing images, and are the corresponding semantic segmentation maps, CM is the binary change map, and SCD1 and SCD2 are the corresponding semantic change detection maps. Figure 5 is the ground truth label corresponding to the binary change map and the semantic change detection map.

[0150] It can be seen from the experimental results that a multi-branch collaborative semantic change detection system and method proposed by the present invention can finely segment ground object categories with relatively small sizes, and the boundaries are relatively clear, the object contours are relatively accurate, the noise content is less, the problem of difficult-to-separate samples and the problem of unbalanced samples are alleviated to a certain extent, the changed and unchanged regions are separated relatively accurately, and the detailed ground object category detection of the changed region has achieved relatively excellent performance.

[0151] The present invention introduces a multi-resolution bidirectional feature fusion strategy during feature extraction; introduces a context information interaction strategy to enhance context information; introduces a channel attention mechanism for deep feature fusion; introduces pseudo-labels and weighted cross-entropy loss to alleviate problems such as insufficient data volume, class imbalance, and difficult-to-separate samples; adopts a multi-branch network, and the features of different branches cooperate with each other, embed more object-oriented features for sub-tasks, provide more consistent change information, improve feature utilization rate, and jointly improve the performance of the detection task.

[0152] The present invention makes the following plan for technology iteration:

[0153] 1. Iteration of the basic network

[0154] To achieve high performance of the model, the feature extraction module selects the ResNet34 network that is most suitable for the task as the basic network, and makes a series of improvements on it to make it perform better in the semantic change detection task. When a network with better performance and higher efficiency is proposed, the technology can be updated by using a better basic network.

[0155] 2. Iteration of the improvement module

[0156] The feature extraction network adopts the method of multi-resolution fusion and context information interaction as the basic implementation method of feature extraction, and uses a fusion unit based on the attention mechanism to solve the difficulties of existing semantic change detection algorithms to a certain extent. The feature extraction can be completed by updating the specific implementation methods of multi-resolution fusion and context information interaction in the feature extraction network or from a better perspective. When a better attention module is proposed, use a better attention module to replace the channel attention module currently used in the present invention to further improve the performance of the model and achieve the iteration of the technology.

[0157] 3. Iteration of downstream tasks

[0158] The feature extraction network trained through multi-resolution fusion and context information interaction is used for the semantic change detection task. Since there may be similar problems in other fields, this network can also be used in tasks that require feature extraction such as semantic segmentation, object detection, and image classification.

Claims

1. A multi-branch collaborative semantic change detection method, characterized in that Specifically, it includes the following steps: S1. Divide the semantic change detection sample set, and randomly divide the publicly available semantic change detection data set into a training set and a test set; S2. Read the images in the training set divided in step S1, and perform data processing and data augmentation on them: S3. Use a multi-resolution convolutional neural bi-directional fusion network to extract features from the dual-temporal remote sensing images. The extracted multi-resolution features are X ms , Add an additional high-resolution branch on the basis of the classification backbone network ResNet. The ResNet network is divided into four stages when performing downsampling. The output result of each stage halves the size of the feature map and doubles the number of channels, gradually extracting high-level semantic features. At different stages, "high-low fusion" and "low-high fusion" are respectively performed on the features of the high-resolution branch and the low-resolution branch. Through bidirectional feature fusion, multi-resolution spatial information and semantic information are fully utilized; S4. Introduce a class self-attention mechanism to build a context information interaction module, and use the context information interaction module to enhance the context information. Finally, the enhanced feature after context information interaction is X ce , S401. Perform convolution and dimensionality conversion operations on the multi-resolution feature X extracted in step S3 ms ; S402. Perform an operation on X ms Obtain a rough semantic segmentation result RSS through a simple rough segmentation network; S403. Perform Softmax and dimensionality conversion operations on the rough semantic segmentation result RSS obtained in step S402 to obtain C groups, where C is the segmentation category, and the vector set V = {V i , i = 1,..., C}, and each vector V i is a feature representation of a category; S404. Multiply the results obtained in step S401 and step S403 to obtain the object context representation; S405. Calculate the relationship between each pixel and the object context representation obtained in step S404 to obtain the final representation of each pixel; S5. Channel fusion is performed on the features extracted in the dual time phases, and the finally fused feature is X fus , S6. According to the finally fused features in steps S4 and S5, obtain the final output result through the corresponding semantic segmentation decoder and binary change detection decoder.

2. The multi-branch collaborative semantic change detection method according to claim 1, wherein The specific method of step S2 is as follows: First, select multiple classic network models, including ResNet101 and HRNet-W64. Input the original data into the model for training, and use the trained model to predict the specific semantic categories of the unchanged regions of the semantic change detection data set through model integration. Take the predicted results of the unchanged regions as pseudo-labels, and combine them with the existing labels of the changed regions to achieve secondary annotation of the data set and increase the number of labeled pixels; Then, perform data augmentation on the data by randomly flipping left and right, flipping up and down, and rotating at multiple angles.

3. A multi-branch collaborative semantic change detection method according to claim 1, characterized in that The model used for pseudo-label generation in step S2 does not participate in subsequent training and testing.

4. A multi-branch collaborative semantic change detection method according to claim 1, characterized in that The specific method of step S5 is as follows: S501. For the enhanced features extracted in step S4 perform channel fusion in multiple ways, including three ways: feature addition, subtraction, and splicing, and provide the features obtained by combining multiple ways for the binary change detection network branch, as shown in formulas (11)-(13): S502. Introduce a channel attention mechanism to process the feature channels. Formulas (11)-(13) are along the channel dimension, and different features are obtained through three ways of feature addition, subtraction, and splicing. Perform the following operations on the obtained features: 1) Squeeze operation: Through global average pooling, perform feature compression along the spatial dimension, turning each two-dimensional feature channel into a real number. This real number has a global receptive field to some extent, and the output dimension matches the number of input feature channels; 2) Excitation operation: Generate weights for each feature channel through parameters; 3) Reweighting operation: Regard the weights output by the excitation operation as the importance of each feature channel after feature selection, and then multiply each channel by weight to all the features before step 3) of step S502 through multiplication, completing the recalibration of the original features in the channel dimension. The introduced channel attention mechanism selects more important features through the above operations, and then performs a dimensionality reduction operation to compromise the richness of the features and the computational complexity; X fus = SA(X con ) (14) Among them, SA(·) refers to the channel attention operation; The finally fused feature is X fus , 5. The multi-branch collaborative semantic change detection method according to claim 2, characterized in that, The specific method of step S6 is as follows: The features extracted in steps S4 and S5 are passed through the corresponding semantic segmentation decoder and the binary change detection decoder to obtain the final output results: a binary change map CM, and two semantic change detection maps SCD1 and SCD2; Among them, Mask(·) is a dot product operation. After dot multiplying the semantic segmentation result with the binary change detection map CM, only the semantic information of the changed area is retained; the semantic segmentation decoder and the binary change detection decoder have the structure of a simplified fully convolutional network. Through convolution, BN, and Relu, the segmentation and change results are obtained, and then interpolation is performed to obtain a semantic segmentation map and a binary change map of the same size as the input.

6. A multi-branch collaborative semantic change detection system based on the detection method according to any one of claims 1 to 5, characterized in that, It includes: The multi-resolution convolutional neural network encoder is composed of two identical twin structures with shared weights, which are used to input the dual-temporal images I1 and I2 and extract multi-resolution depth features; The context information interaction module is composed of two identical twin structures with shared weights, which are used to perform context information interaction on the target region and concatenate it with the original features to obtain the final enhanced features; The channel fusion module is used to explore the correlation of hidden deep features, enhance the collaborative interaction between subtasks, provide richer and more effective features for the binary change detection network branch, and improve the feature utilization rate; The semantic segmentation decoder is a twin structure with two weight-sharing and identical structures, and is used to output two semantic segmentation maps and The binary change detection decoder is used to output a binary change map CM, as well as two semantic change detection maps SCD1 and SCD2.

7. A multi-branch collaborative semantic change detection device based on the detection method according to any one of claims 1 to 5, characterized in that It includes: A memory for storing computer programs, data, and models; A processor for implementing the multi-branch collaborative semantic change detection method described in steps S1 to S6 when executing the computer program.

8. A computer-readable storage medium, characterized in that, Read and store programs and data. The computer-readable storage medium stores a computer program, which can perform multi-branch collaborative semantic detection based on the detection method described in any one of claims 1 to 5 when executed by the processor.

Citation Information

Patent Citations

  • Real-time street view image semantic segmentation method based on deep multi-branch aggregation

    CN113011336A

  • Remote sensing image change detection method and system fusing regional semantics and pixel features

    CN115331087A