A dual-path remote sensing image change detection method, system and device based on pixel and semantic information interaction

CN118015479BActive Publication Date: 2026-09-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410083833.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2026-09-29
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

[0013]2)在涉及变化检测时,很少有专门研究通道维度信息交互的方法

Benefits of technology

[0103]1)本发明通过设计一种基于像素和语义信息交互的双路径遥感影像变化检测方法,该方法包括主干特征提取网络、像素语义双路径提取交互模块和变化检测头,其中像素语义双路径提取交互模块包含:通道打乱卷积模块、Transformer编码器和像素语义交互模块。有效促进不同时相数据在通道维度的信息交互,能够提取更为丰富的差异信息,从而协助模型更轻松、准确地捕捉水体变化的范围。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118015479B_ABST
    Figure CN118015479B_ABST
Patent Text Reader

Abstract

A kind of dual-path remote sensing image change detection method, system and equipment based on pixel and semantic information interaction, method includes: obtaining remote sensing image change detection dataset, and its data annotation and image block, obtain CDDS dataset;CDDS dataset is preprocessed and divided into training, verification and test set, and data enhancement is carried out to training and verification set;Dual-path change detection network PST-Net based on pixel and semantic information interaction is constructed;Different hyperparameter schemes are designed for PST-Net network;PST-Net network is trained using training set and verification set, and the optimal model under multiple different hyperparameter schemes is obtained;The optimal model under multiple different hyperparameter schemes obtained is tested using training set and verification set, and the optimal model under the optimal hyperparameter scheme is obtained;The optimal model is tested using test set, and multiple evaluation indexes are obtained;System and equipment are used to realize the method;The present application has the advantages of high recognition rate, low computing cost, fast recognition speed and modular design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image change detection technology, specifically relating to a dual-path remote sensing image change detection method, system, and device based on pixel and semantic information interaction. Background Technology

[0002] Change detection technology is an important branch of remote sensing. By analyzing and comparing remote sensing images at different times, it identifies changes in the state of ground features, providing strong support for national resource updates, disaster assessment, and urban planning. In recent years, with the rapid development of SAR technology and artificial intelligence, SAR image change detection has gained more possibilities in data acquisition and algorithm processing, becoming a research hotspot.

[0003] Change detection technology, as a key technology, is not only of great value in flood disaster monitoring, but also plays a crucial role in land use, urban planning, resource management, and post-disaster reconstruction. By analyzing and comparing multi-temporal remote sensing images, change detection technology can effectively identify changes in the land surface or scene, providing support for government decision-making, disaster assessment, and resource management.

[0004] Accurate detection of flood disaster changes is crucial for disaster relief and mitigation, making remote sensing and change detection technologies particularly important. Remote sensing mainly consists of optical images and Synthetic Aperture Radar (SAR) images. However, optical images are easily affected by weather and fog, leading to interference in detection. Since the advent of SAR images, more researchers have favored using SAR images to identify water bodies. While SAR images offer advantages such as all-weather, all-time availability, and immunity to fog and lighting conditions, they are a better choice for water area segmentation. In recent years, the resolution of SAR images has been continuously improving, currently reaching sub-decimeter levels. This provides more possibilities for achieving high-precision change detection and supports the rapid identification of disaster-stricken areas and the reduction of disaster losses.

[0005] The basic workflow for SAR image change detection includes: 1) image preprocessing; 2) difference map generation; and 3) difference map analysis. First, SAR images from two different time periods are preprocessed, including image registration, geometric correction, radiometric correction, and image filtering. Image registration ensures that pixels at the same geographical location correspond consistently in both images, thus guaranteeing the accuracy of subsequent evaluation metrics. Geometric correction reduces geometric distortion, while radiometric correction reduces radiometric distortion. Image filtering helps reduce speckle noise, which is beneficial for subsequent difference map generation and analysis. Second, specific methods are used to identify the differences between corresponding pixels in the two preprocessed SAR images, presenting them as a difference map. In the difference map, pixels with higher grayscale values ​​indicate a greater likelihood of change, while pixels with lower grayscale values ​​indicate a lower likelihood of change. The final detection accuracy is significantly affected by the quality of the difference map. The final step is to perform pixel-by-pixel binary classification of the difference map, dividing it into changed and unchanged classes. Traditional difference map analysis methods include thresholding, clustering, and some machine learning methods. These methods generally have fast computation speeds. In addition, although deep learning-based methods take longer to compute, they are able to learn more image features.

[0006] Existing methods for detecting changes in remote sensing images based on supervised learning can be broadly divided into two main schools of thought.

[0007] The first approach is a two-stage method, which trains two temporal images using either a Convolutional Neural Network (CNN) or a Fully Convolutional Network (FCN), and then compares the classification results to derive the change detection result. This method can only be implemented when two change labels and two temporal semantic labels are available.

[0008] Another approach is the single-stage method, which directly extracts change results from bi-temporal images. Block-level methods model change detection as a similarity detection process, dividing the bi-temporal images into a pair of blocks and using a CNN to obtain the center prediction of this pair of image blocks. Pixel-level methods, on the other hand, use an FCN to directly extract high-resolution change maps from the two inputs, which is generally more effective and efficient than block-level methods. Since change detection involves processing two inputs, how to fuse bi-temporal information is an important issue. Existing FCN-based methods can be roughly divided into two groups based on the stage of bi-temporal information fusion. Image-level methods feed the stitched bi-temporal images as a single input into a semantic segmentation network. Feature-level methods fuse the bi-temporal features extracted by the neural network, making change decisions based on the fused features.

[0009] Another approach is the single-stage method, which directly derives change results from bi-temporal images. Block-level methods model change detection as a similarity detection process by dividing the bi-temporal images into pairs of blocks and then using a CNN to predict the center of these blocks. Pixel-level methods use FCNs to directly obtain a high-resolution change map from the two inputs, which is generally more efficient and effective than block-level methods. Since change detection requires solving two inputs, how to fuse bi-temporal information is a crucial issue. Existing FCN-based methods can be roughly divided into two groups based on the stage of bi-temporal information fusion. Image-level methods stitch the bi-temporal images together as a single input and feed them into a semantic segmentation network. Feature-level methods fuse bi-temporal features extracted from the neural network, making change decisions based on the fused features.

[0010] OmbriaNet (DRAKONAKIS GI, TSAGKATAKIS G, FOTIADOU K, et al. OmbriaNet—Supervised Flood Mapping via Convolutional Neural Networks UsingMultitemporal Sentinel-1and Sentinel-2Data Fusion[J / OL]. IEEE Journal ofSelected Topics in Applied Earth Observations and Remote Sensing,2022,15:2341-2356.DOI:10.1109 / JSTARS.2022.3155559.) is a semantic segmentation network U-Net (RONNEBERGERO, FISCHER P, BROX TU-Net: Convolutional Networks for Biomedical The U-Net network, an improved version of ImageSegmentation [M / OL].arXiv, 2015 [2023-12-15]. http: / / arxiv.org / abs / 1505.04597., consists of two parts: a downsampling path and an upsampling path. In the downsampling path, a series of convolutional and pooling layers downsample the input image into small feature maps. In the upsampling path, a series of convolutional and upsampling layers restore the small feature maps to their original size, yielding pixel-level predictions. Furthermore, U-Net employs skip connections, linking feature maps from the downsampling path with those from the corresponding upsampling path, allowing the network to utilize both low-level and high-level features for prediction. OmbriaNet improves upon the encoder-decoder architecture by creating meaningful feature maps from multi-temporal images, enhancing change detection accuracy. This network detects changes based on feature maps generated from temporal images. OmbriaNet uses regions from two different timestamps in a dual-temporal image as input: the first image before the event and the second image after the event.

[0011] In summary, existing technologies rarely focus on methods for detecting changes in remotely sensed images of water bodies; instead, they primarily study changes in buildings (CHEN J, YUAN Z, PENG J, et al. DASNet: Dual attentive fully convolutional siamese networks for change detection of high resolution satellite images[J / OL]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14:1194-1206. DOI:10.1109 / JSTARS.2020.3037893.). Furthermore, most existing algorithms for detecting changes in water bodies are based on convolutional networks, with few combining them with Transformers. In conclusion, existing technologies have the following drawbacks:

[0012] 1) Currently, research on change detection based on remote sensing images of water areas is relatively limited. Most studies focus more on the change detection of objects such as buildings.

[0013] 2) When it comes to change detection, there are very few methods specifically studied for the interaction of channel-dimensional information.

[0014] 3) Current methods mainly process multi-temporal input data separately, lacking a comprehensive consideration of the interaction between pixel-level information and high-level semantic information.

[0015] 4) Existing change detection algorithms are highly complex, computationally intensive, and lack lightweight features. Summary of the Invention

[0016] To overcome the shortcomings of the prior art, the present invention aims to provide a dual-path remote sensing image change detection method, system, and device based on pixel and semantic information interaction. Combining pixel-level and semantic-level information helps to improve the accuracy and reliability of change detection. Pixel-level information can capture subtle changes, while semantic information helps to understand the meaning and context of changes, reducing false alarms and false negatives. It makes full use of multi-scale information of the image, improves the accuracy, robustness, and application range of change detection, and makes the detection results more reliable and meaningful.

[0017] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0018] A dual-path remote sensing image change detection method based on pixel and semantic information interaction includes the following steps:

[0019] Step 1: Obtain the remote sensing image change detection dataset, and perform data annotation and image patching to obtain the CDDS dataset;

[0020] Step 2: Preprocess the CDDS dataset obtained in Step 1, divide the preprocessed CDDS dataset into training set, validation set and test set, and perform data augmentation on the training set and validation set;

[0021] Step 3: Construct the PST-Net dual-path change detection network based on pixel and semantic information interaction;

[0022] Step 4: Design multiple sets of different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction constructed in Step 3;

[0023] Step 5: Based on the multiple sets of different hyperparameter schemes designed in Step 4, use the training set and validation set processed in Step 2 to train the dual-path change detection network PST-Net built in Step 3 based on pixel and semantic information interaction, and obtain the optimal model under multiple sets of different hyperparameter schemes.

[0024] Step 6: Use the training set and validation set processed in Step 2 to train the optimal models under multiple different hyperparameter schemes obtained in Step 5, and obtain the optimal model under the optimal hyperparameter scheme.

[0025] Step 7: Using the optimal model obtained in Step 6, test it using the test set divided in Step 2 to obtain various evaluation metrics.

[0026] The specific method for step 1 is as follows:

[0027] Step 1.1: Collect remote sensing images of the same area at two different time points;

[0028] Step 1.2: Data Labeling:

[0029] Use the Labelme tool to annotate the remote sensing images obtained in step 1.1, mark the areas that have changed, and create labels;

[0030] Step 1.3: Image segmentation:

[0031] The remote sensing images obtained in step 1.1 and the labels obtained in step 1.2 are divided into blocks, and the blocks and corresponding labels are combined to form a dataset, namely the CDDS dataset.

[0032] The specific method for step 2 is as follows:

[0033] Step 2.1: Normalize the CDDS dataset obtained in Step 1.3:

[0034] Calculate the overall maximum and minimum values ​​for the same channel across all images, following the channel dimension. Then, normalize the data for that channel using these overall maximum and minimum values, using the following formula:

[0035]

[0036]

[0037] in, This represents the value of the j-th pixel in the c-th channel of the i-th image. express Normalized pixel value, min c The minimum pixel value of the c-th channel layer, max c The maximum pixel value of the c-th channel layer;

[0038] Step 2.2: After normalization, layers with the same channel in all images are grouped together. Within each group, the population mean and population variance of the data obtained in Step 2.1 are calculated. The formulas for calculating the population mean and population variance are as follows:

[0039]

[0040]

[0041] Among them, Mean c Std represents the pixel mean of the c-th channel. c This represents the standard deviation of pixels in the c-th channel. Let N represent the value of the j-th pixel in the c-th channel of the i-th image, N represent the number of images, and M represent the number of pixels in a single image.

[0042] Step 2.3: Data Standardization Processing

[0043] For all the normalized data obtained in step 2.2 Standardization is performed at the channel level, using the following formula:

[0044]

[0045] in, This represents the standardized value of the j-th pixel in the c-th channel of the i-th image;

[0046] Step 2.4: Divide the standardized data from Step 2.3 into training, validation, and test sets;

[0047] Step 2.5: Perform data augmentation on the training and validation sets from Step 2.4:

[0048] We use "vertical flipping + horizontal flipping", "shifting", and "scaling" to augment the training and validation sets.

[0049] The specific method for step 3 is as follows:

[0050] The dual-path change detection network PST-Net based on pixel and semantic information interaction includes: a backbone feature extraction network, a pixel-semantic dual-path extraction and interaction module, and a change detection head;

[0051] Step 3.1: Construct the backbone feature extraction network:

[0052] The backbone feature extraction network is ResNet18. First, the features of phase 1 and phase 2 are passed through the ResNet18 network. After two downsamplings, the features become 1 / 4 of the original. After two more downsamplings, the features become 1 / 16 of the original. The 1 / 4 and 1 / 16 features are then input into the pixel semantic dual-path extraction interaction module for further processing.

[0053] Step 3.2: Construct a pixel-semantic dual-path extraction interaction module:

[0054] The network consists of three parts: Channel Shuffle Conv Block (CSCB), Transformer Encoder, and Pixel and Semantic Interaction Model (PSIM).

[0055] Step 3.2.1: Construct the CSCB channel shuffling convolution module:

[0056] The CSCB channel shuffling convolution module consists of dual-temporal pixel-level feature concatenation, pooling, channel convolution (CCM), and upsampling.

[0057] The CCM channel convolution module performs channel convolution in the C×HW dimension; the specific operation is as follows:

[0058] 1) B×C×H×W=>B×C×HW, first straighten the feature map in the spatial dimension (HW);

[0059] 2) B×C×HW=>B×NC×C×HW, adding a channel dimension, NC, to the second dimension. NC refers to the new channel. After the transformation, a matrix that conforms to the convolution operation is constructed. The initial value of NC is 1.

[0060] 3) Shuffle: 3×3 convolution is used for sliding to interact information between different channels. Before performing convolution, a channel shuffling operation is performed.

[0061] 4) After performing convolution, perform inverse shuffle and B×1×C×HW => B×C×H×W;

[0062] Step 3.2.2: Construct a multi-layer Transformer encoder;

[0063] The Transformer encoder consists of Patch Embedding, LayerNorm, Multi-Head Attention (MHA), and Multilayer Perceptron (MLP).

[0064] The specific process is as follows: First, the image is embedded into blocks to obtain an image block sequence, where each image block is represented as a vector; then, the image block sequence is input to the normalization layer, followed by multi-head attention calculation and residual connection; finally, after layer normalization and multi-layer perceptron, the output is generated; the Transformer encoder will be stacked multiple times to form a multi-layer Transformer encoder.

[0065] Step 3.2.3: Construct the PSIM pixel semantic interaction module:

[0066] The PSIM pixel semantic interaction module overlaps N layers of multi-head attention. The structure includes 4 LayerNorms, 2 Feed Forward Networks (FFNs), and 2 multi-head self-attentions. This module has two inputs, which are initially features from the pixel path and the semantic path, respectively. The pixel path is used to extract low-level features of the image, while the semantic path focuses on high-level information in the image.

[0067] Step 3.3: Construct the change detection head;

[0068] At this point, the pixel features of phase 1 and phase 2, after N layers of interaction in the PSIM pixel semantic interaction module, contain rich semantic information. After subtraction, the difference map of the two phases is obtained. The change detection head then performs two 3×3 convolutions on the obtained difference map and directly upsamples it by 4 times to become a change map with the same size as the original input, which is the prediction result map.

[0069] The specific method for step 4 is as follows:

[0070] Step 4.1: For the four hyperparameters Layer, STL (Semantictokenlength), Dim and KS (Kernel Size) in the dual-path change detection network PST-Net based on pixel and semantic information interaction constructed in Step 3, design multiple sets of different values ​​for each.

[0071] Step 4.2: Using the control variable method, multiple sets of values ​​for the four hyperparameters in Step 4.1 are combined to obtain multiple different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction.

[0072] The specific method for step 5 is as follows:

[0073] Using the training and validation sets in the CDDS dataset processed in step 2, we trained and tested multiple hyperparameter schemes of the dual-path change detection network PST-Net obtained in step 4.2 based on pixel and semantic information interaction. We analyzed the test results and finally obtained the optimal hyperparameter scheme through model evaluation.

[0074] The specific method for step 6 is as follows:

[0075] Step 6.1: Construct the dual-path change detection network PST-Net based on pixel and semantic information interaction in Step 3 according to the optimal hyperparameter scheme obtained in Step 5.2;

[0076] Step 6.2: Train the PST-Net network constructed in Step 6.1 using the training and validation sets processed in Step 2; during training, the loss function used is the Binary Cross Entropy (BCE) loss function L. log (y, p(x)) is defined as:

[0077] L log (y,p(x))=-(ylog(p(x))+(1-y)log(1-p(x))) Equation 6

[0078] Where y is the true label and p(x) is the predicted probability of the model;

[0079] Step 6.3: During each training process, save the model with the highest F1 coefficient on the validation set after Step 2 as the optimal model.

[0080] The specific method for step 7 is as follows:

[0081] Based on the optimal model with the highest F1 coefficient obtained in step 6, test it using the test set from step 2. Use the F1 score, pixel accuracy (PA), mean intersection over union (mIoU), frequency weighted intersection over union (FWIoU), and Kappa coefficient as evaluation metrics.

[0082] The formulas for calculating the evaluation metrics F1 (F1-score), Pixel Accuracy (PA), Mean Intersection over Union (mIoU), Frequency Weighted Intersection over Union (FWIoU), and Kappa coefficient are as follows:

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] The meanings of the parameters in equations 7, 8, 9, 10, 11, and 12 are as follows:

[0090] TP (True Positive): Predicted correctly, the predicted result is positive, and the actual value is positive. FP (False Positive): Predicted incorrectly, the predicted result is positive, and the actual value is negative. FN (False Negative): Predicted incorrectly, the predicted result is negative, and the actual value is positive. TN (True Negative): Predicted correctly, the predicted result is negative, and the actual value is negative.

[0091] This invention also provides a dual-path remote sensing image change detection system based on pixel and semantic information interaction, comprising:

[0092] Dataset building module: used to acquire remote sensing image change detection dataset, and to perform data annotation and image patching processing to obtain CDDS dataset;

[0093] Dataset processing module: Used to preprocess the CDDS dataset, divide the preprocessed CDDS dataset into training set, validation set and test set, and perform data augmentation on the training set and validation set;

[0094] Dual-path change detection network building module: used to build the dual-path change detection network PST-Net based on pixel and semantic information interaction;

[0095] Multiple hyperparameter scheme design module: used to design multiple different hyperparameter schemes for the constructed dual-path change detection network PST-Net based on pixel and semantic information interaction;

[0096] The module for obtaining the optimal model under multiple different hyperparameter schemes is used to train the constructed dual-path change detection network PST-Net based on pixel and semantic information interaction using the processed training and validation sets according to multiple different hyperparameter schemes, so as to obtain the optimal model under multiple different hyperparameter schemes.

[0097] The module for obtaining the optimal model under the optimal hyperparameter scheme is used to train multiple sets of optimal models under different hyperparameter schemes using the processed training set and validation set, and obtain the optimal model under the optimal hyperparameter scheme.

[0098] Multiple evaluation metrics acquisition module: Used to obtain multiple evaluation metrics by using the optimal model and the partitioned test set for testing.

[0099] The present invention also provides a dual-path remote sensing image change detection device based on pixel and semantic information interaction, comprising:

[0100] Memory: A computer program that stores the above-mentioned dual-path remote sensing image change detection method based on pixel and semantic information interaction, and is a computer-readable device;

[0101] Processor: Used to implement the dual-path remote sensing image change detection method based on pixel and semantic information interaction when executing the computer program.

[0102] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0103] 1) This invention designs a dual-path remote sensing image change detection method based on pixel and semantic information interaction. This method includes a backbone feature extraction network, a pixel-semantic dual-path extraction and interaction module, and a change detection head. The pixel-semantic dual-path extraction and interaction module comprises a channel shuffling convolution module, a Transformer encoder, and a pixel-semantic interaction module. This effectively promotes information interaction between data from different time phases along the channel dimension, enabling the extraction of richer difference information, thereby assisting the model in more easily and accurately capturing the range of water body changes.

[0104] 2) This invention makes full use of the backbone feature extraction network to extract dual-temporal features at different scales, and designs a method for interaction between pixel information and semantic information, which can extract and fuse change features at multiple scales, helping the model to perceive water changes more accurately.

[0105] 3) In the CSCB module constructed in step 3.2.1 of this invention, pooling operation is used before the CCM module to reduce the resolution of the feature map, effectively reducing the model size and improving computational efficiency. Therefore, compared with the current change detection technology, it has a lower number of parameters and computational load, ensuring that the model can run more efficiently even on devices with lower performance.

[0106] 4) The model in this invention adopts a modular design approach. The backbone feature extraction network, the pixel semantic dual-path extraction interaction module, and the basic network structure of the detection head can all be replaced. Modules can be added or modified according to the shortcomings of the basic network in the prediction task. With the development of emerging technologies and the proposal of better network modules, this invention can be iteratively updated at any time to improve model performance.

[0107] In summary, this invention, by designing a method for channel information fusion and multi-scale information interaction, effectively helps the model to perceive water changes more accurately, and has the advantages of high recognition rate, low computational cost, fast recognition speed and modular design. Attached Figure Description

[0108] Figure 1 This is a diagram of the overall structure of the dual-path change detection network for pixel and semantic information interaction in this invention.

[0109] Figure 2 This is a structural diagram of the Channel Shuffle Conv Block (CSCB) of this invention.

[0110] Figure 3 This is a structural diagram of the Pixel and Semantic Interaction Model (PSIM) of the present invention.

[0111] Figure 4 This is the overall flowchart of the present invention.

[0112] Figure 5 Training flowchart of this invention.

[0113] Figure 6 Flowchart of the test process for this invention.

[0114] Figure 7 This is a visualization of the experimental results of this invention. Detailed Implementation

[0115] The technical solution adopted by the present invention will be further described below with reference to the accompanying drawings.

[0116] This invention proposes a dual-path change detection network based on pixel and semantic information interaction, utilizing SAR image remote sensing data. The method aims to design a dual-path system: introducing a pixel path and a semantic path to process low-level and high-level information respectively. Through interactive learning, it effectively aggregates multi-scale pixel features and semantic information, improving the sensitivity and accuracy of image change detection. An information fusion and interaction module is introduced: a PSIM module is incorporated to promote the interaction and fusion between pixel and semantic information. This information interaction enhances the model's understanding of information at different levels, improving detection robustness. A CSCB (Channel Convolution) module is introduced: using the CCM module for convolution operations along the channel dimension, it more accurately captures feature differences between image channels, thereby improving the accuracy and sensitivity of change detection. The network as a whole is a lightweight structure, reducing the number of model parameters and computational complexity while maintaining accuracy and enabling rapid training and deployment, making it more suitable for practical applications. These innovations result in higher accuracy, sensitivity, and practicality in the field of remote sensing image change detection, providing a novel and more effective method for processing changes in multi-temporal images.

[0117] The overall flowchart of this invention is shown below. Figure 4 A dual-path remote sensing image change detection method based on pixel and semantic information interaction includes the following steps:

[0118] Step 1: Obtain the remote sensing image change detection dataset, and perform data annotation and image patch processing on it to obtain the CDDS (Change Detection DataSet) dataset;

[0119] The specific method for step 1 is as follows:

[0120] Step 1.1: Collect remote sensing images of the same area before and after the flood;

[0121] Step 1.2: Data Labeling:

[0122] Use the Labelme tool to annotate the remote sensing images obtained in step 1.1, mark the areas that have changed, and create labels;

[0123] Step 1.3: Image segmentation:

[0124] The remote sensing images obtained in step 1.1 and the labels obtained in step 1.2 are processed into blocks. Specifically, the remote sensing images and labels are divided into blocks of 256×256 pixels, with an overlap of 32 pixels between adjacent images. The blocks and the corresponding labels are combined to form a dataset called the CDDS dataset.

[0125] Step 2: Preprocess the CDDS dataset obtained in Step 1, divide the preprocessed CDDS dataset into training set, validation set and test set, and perform data augmentation on the training set and validation set;

[0126] The specific method for step 2 is as follows:

[0127] Step 2.1: Normalize the CDDS dataset obtained in Step 1.3:

[0128] Calculate the overall maximum and minimum values ​​for the same channel across all images, following the channel dimension. Then, normalize the data for that channel using these overall maximum and minimum values, using the following formula:

[0129]

[0130]

[0131] in, This represents the value of the j-th pixel in the c-th channel of the i-th image. express Normalized pixel value, min c The minimum pixel value of the c-th channel layer, max c The maximum pixel value of the c-th channel layer;

[0132] Step 2.2: After normalization, layers with the same channel in all images are grouped together. Within each group, the population mean and population variance of the data obtained in Step 2.1 are calculated. The formulas for calculating the population mean and population variance are as follows:

[0133]

[0134]

[0135] Among them, Mean c Std represents the pixel mean of the c-th channel. c This represents the standard deviation of pixels in the c-th channel. Let N represent the value of the j-th pixel in the c-th channel of the i-th image, N represent the number of images, and M represent the number of pixels in a single image.

[0136] Step 2.3: Data Standardization Processing

[0137] The normalized data X obtained in step 2.2 is then standardized along the channel dimension. The standardization formula is as follows:

[0138]

[0139] in, This represents the standardized value of the j-th pixel in the c-th channel of the i-th image;

[0140] Step 2.4: Divide the standardized data from Step 2.3 into training, validation, and test sets in a 6:2:2 ratio;

[0141] Step 2.5: Perform data augmentation on the training and validation sets from Step 2.4;

[0142] Data augmentation was performed on the training and validation sets using techniques such as vertical and horizontal flipping, displacement, and scaling. Since precisely labeled data is scarce in change detection tasks, data augmentation generates more training and validation data, helping to make the data as close as possible to the true distribution and thus improving the model's generalization ability. It's important to note that to ensure consistency between the unaffected and post-change images during data augmentation, this step applies to both the training and validation images simultaneously. After data augmentation, both the training and validation data were expanded to four times their original size.

[0143] like Figure 1 As shown, step 3: Construct the dual-path change detection network PST-Net based on pixel and semantic information interaction;

[0144] This invention designs a change detection network for pixel-level and semantic-level information interaction, which combines several innovations, including dual-path feature extraction, pixel-semantic information fusion, dual-temporal channel information interaction, and a lightweight network. The overall network structure diagram is shown below. Figure 1 Through these innovations, this model is expected to provide a more accurate and efficient solution in the field of remote sensing image change detection. This invention comprises three parts;

[0145] The specific method for step 3 is as follows:

[0146] The dual-path change detection network PST-Net based on pixel and semantic information interaction includes: a backbone feature extraction network, a pixel-semantic dual-path extraction and interaction module, and a change detection head;

[0147] Step 3.1: Construct the backbone feature extraction network:

[0148] The backbone feature extraction network is ResNet18. Specifically, phase 1 and phase 2 are first passed through the ResNet18 network. After two downsampling operations, the features become 1 / 4 of the original. After two more downsampling operations, the features become 1 / 16 of the original. The 1 / 4 and 1 / 16 features are then input into the pixel semantic dual-path extraction interaction module for further processing.

[0149] Step 3.2: Construct a pixel-semantic dual-path extraction interaction module:

[0150] The network consists of three parts: a Channel Shuffle Conv Block (CSCB), a Transformer Encoder, and a Pixel and Semantic Interaction Model (PSIM).

[0151] like Figure 2 As shown, step 3.2.1: Construct a CSCB channel shuffling convolution module;

[0152] The purpose of change detection is to find differences between different images; therefore, the designed model should focus more on channel-level relationships. Based on this consideration, this invention designs a Channel Shuffled Convolutional Module (CSCB), as detailed below. Figure 2 This module consists of dual-temporal pixel-level feature concatenation, pooling, channel convolution (CCM), and upsampling.

[0153] The CCM channel convolution module differs from ordinary convolution, which operates in the H×W dimension; channel convolution operates in the C×HW dimension. This is because, whether for classification, segmentation, or object detection, H×W-dimensional convolution often focuses on the relationships between pixels within the same image. Based on this, the present invention designs a CCM module, the specific operation of which is as follows:

[0154] 1) B×C×H×W=>B×C×HW, first straighten the feature map in the spatial dimension (HW);

[0155] 2) B×C×HW=>B×NC×C×HW. In order to utilize the 2D convolution operator, a channel dimension, namely NC, is added to the second dimension. NC refers to the new channel. After transformation, a matrix that conforms to the convolution operation is constructed. The initial value of NC is 1.

[0156] 3) Shuffle: 3×3 convolution is used for sliding to interact information between different channels. In order to ensure that each window contains data from two time phases, a channel shuffling operation is performed before convolution.

[0157] 4) After performing convolution, perform inverse shuffle and B×1×C×HW => B×C×H×W;

[0158] Step 3.2.2: Construct a multi-layer Transformer encoder;

[0159] The Transformer encoder consists of patch embedding, layer normalization, multi-head attention (MHA), and a multilayer perceptron (MLP), as shown in the structure below. Figure 1 The subgraph is shown.

[0160] The specific process is as follows: First, the image is embedded to obtain an image block sequence, where each image block is represented as a vector; then, the image block sequence is input to the normalization layer, followed by multi-head attention calculation and residual connection; finally, after layer normalization and multi-layer perceptron, the output is generated; the Transformer encoder is stacked multiple times to form a multi-layer Transformer encoder, where each layer can gradually extract and combine high-level feature representations of the image blocks.

[0161] like Figure 3 As shown, step 3.2.3: Construct the Pixel Semantic Interaction Module (PSIM);

[0162] This module incorporates N layers of multi-head attention, and its structure includes 4 LayerNorms, 2 FeedForward Networks (FFNs), and 2 multi-head self-attention networks. See the detailed module description below. Figure 3 This module has two inputs: features from the pixel path and the semantic path, respectively. The pixel path is used to extract low-level features of the image, such as color and texture; the semantic path focuses on high-level information such as objects and scenes in the image. By designing these two paths in parallel, multi-level information fusion and interaction are achieved, improving the model's ability to understand and utilize information at different scales.

[0163] Step 3.3: Construct the change detection head;

[0164] At this point, the pixel features of phase 1 and phase 2, after N layers of interaction in the PSIM pixel semantic interaction module, contain rich semantic information. After subtraction, the difference map of the two phases is obtained. The change detection head then performs two 3×3 convolutions on the obtained difference map and directly upsamples it by 4 times to become a change map with the same size as the original input, which is the prediction result map.

[0165] Step 4: Design multiple sets of different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction constructed in Step 3;

[0166] The specific method for step 4 is as follows:

[0167] Step 4.1 For the four hyperparameters Layer, STL (Semantic token length), Dim, and KS (Kernel Size) in the dual-path change detection network PST-Net constructed in Step 3, multiple sets of different values ​​are designed for each; the Layer hyperparameter represents... Figure 1 The middle part overlaps N times in total, namely the CSCB channel scrambling convolution module, the Transformer encoder, and the Pixel Semantic Interaction Module (PSIM); the STL (Semantic Token Length) hyperparameter is... Figure 1 The number of token sets in the dataset; the Dim hyperparameter is the depth of semantic and pixel features, i.e., the length of each semantic or pixel feature; the KS (Kernel Size) hyperparameter represents... Figure 2 The Conv sizes in the CCM are 3×3, 1×5, 1×7, 1×9, and 1×11, respectively.

[0168] Step 4.2: Using the control variable method, multiple sets of values ​​for the four hyperparameters in Step 4.1 are combined to obtain multiple different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction.

[0169] The following experiments were conducted with four sets of hyperparameters.

[0170] ①Layer = 3, 4, 5, 6;

[0171] ②STL = 4, 6, 8;

[0172] ③Dim = 32, 48, 64, 96;

[0173] ④KS=3×3,1×5,1×7,1×9,1×11;

[0174] Step 5: Based on the multiple sets of different hyperparameter schemes designed in Step 4, use the training set and validation set processed in Step 2 to train the dual-path change detection network PST-Net built in Step 3 based on pixel and semantic information interaction, and obtain the optimal model under multiple sets of different hyperparameter schemes.

[0175] The specific method for step 5 is as follows:

[0176] The training and validation sets in the CDDS dataset processed in step 2 were used to train and test the multiple hyperparameter schemes constructed in step 4.2; the test results were analyzed, and finally, the optimal hyperparameter scheme was obtained through model evaluation; the experimental results are shown in Table 1. The optimal configuration scheme is: Layer=4, STL=4, Dim=64, KS=3×3.

[0177] Table 1 Experimental results for the four hyperparameters

[0178] F1 80.978 82.941 81.983 82.067 STL 4 6 8 F1 82.941 80.175 81.359 Dim 32 48 64 96 F1 81.344 81.243 82.941 81.611 KS 3×3 1×5 1×7 1×9 1×11 F1 82.941 80.589 81.923 81.970 81.969

[0179] Step 6: Use the training set and validation set processed in Step 2 to train the optimal models under multiple different hyperparameter schemes obtained in Step 5, and obtain the optimal model under the optimal hyperparameter scheme.

[0180] The specific method for step 6 is as follows:

[0181] Step 6.1: See details of the training process. Figure 6 Based on the optimal hyperparameter scheme obtained in step 5, the dual-path change detection network PST-Net, which is based on the interaction of pixel and semantic information, is constructed in step 3.

[0182] Step 6.2: Train the PST-Net constructed in Step 6.1 using the CDDS training and validation sets processed in Step 2; during training, the loss function used is the Binary Cross Entropy (BCE) loss function L. log (y, p(x)) is defined as:

[0183] L log (y,p(x))=-(ylog(p(x))+(1-y)log(1-p(x))) Equation 6

[0184] Where y is the true label and p(x) is the predicted probability of the model;

[0185] Step 6.3: During each training process, save the model with the highest F1 coefficient on the CDDS validation set after Step 2 as the optimal model.

[0186] like Figure 6 As shown, step 7: Using the optimal model obtained in step 6, test it using the CDDS test set divided in step 2 to obtain various evaluation metrics.

[0187] The specific method for step 7 is as follows:

[0188] Based on the optimal model with the highest F1 coefficient obtained in step 6, the model is tested using the CDDS test set from step 2. The model's F1 score, pixel accuracy (PA), mean intersection over union (mIoU), frequency weighted intersection over union (FWIoU), and Kappa coefficient are used as evaluation metrics.

[0189] The formulas for calculating the evaluation metrics F1 (F1-score), Pixel Accuracy (PA), Mean Intersection over Union (mIoU), Frequency Weighted Intersection over Union (FWIoU), and Kappa coefficient are as follows:

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196] The meanings of the parameters in equations 7, 8, 9, 10, 11, and 12 are as follows:

[0197] TP (True Positive): Predicted correctly, the predicted result is positive, and the actual value is positive. FP (False Positive): Predicted incorrectly, the predicted result is positive, and the actual value is negative. FN (False Negative): Predicted incorrectly, the predicted result is negative, and the actual value is positive. TN (True Negative): Predicted correctly, the predicted result is negative, and the actual value is negative.

[0198] This invention designs a dual-path interaction module to fully integrate low-level pixel information and high-level semantic information, thereby enhancing the extraction capability of ground feature characteristics. Pixel path features are rich in texture information, while semantic path features contain more aggregate information and can perceive a wider range of features. Based on this concept, this invention processes the features extracted by the backbone network through both pixel path and semantic path processing to obtain ground feature description information at different levels. Subsequently, by cleverly designing an interaction module between the pixel path and semantic path, it effectively promotes full interaction between high-level and low-level features. To achieve lightweight model and efficient interaction of channel dimension information, this invention uses a convolutional neural network (CNN) as the feature extractor for the pixel path and introduces a CCM module designed with channel shuffling operations. This module significantly improves the accuracy of change detection by leveraging channel convolution technology. This innovation fully combines high- and low-level information interaction with channel dimension information interaction, providing an effective solution for improving the performance of ground feature change detection.

[0199] In summary, this invention is not limited to flood disaster monitoring, but also has broad application prospects in various fields, providing an efficient and reliable change detection solution for responding to natural disasters and resource management.

[0200] This invention aims to address the challenges of change detection in remote sensing images by combining pixel-level and semantic-level information to improve the accuracy of change detection. Pixel-level information helps capture minute changes in images, thereby improving the performance of change detection.

[0201] As can be seen from the embodiments, compared with the prior art, the present invention has the following innovative points:

[0202] 1. Iteration of ordinary channel modules

[0203] To achieve both efficiency and high performance, this invention selects ResNet18, the most suitable network for this task, as the backbone feature extraction network. It can then be combined with depthwise separable convolutions to further improve the accuracy of change detection. When networks with better performance and higher efficiency are proposed, the technology can be updated by using superior feature extraction networks.

[0204] 2. Iterative approach using a single-branch method for routine change detection.

[0205] The encoder of this invention employs a feature extraction structure with dual-temporal interleaved channel information interaction to extract change information. This avoids interference that occurs when directly splicing dual temporal phases during feature extraction to a certain extent, thus improving the performance of water area change detection. Therefore, this invention can achieve technological iteration by updating the specific implementation methods in the change detection framework. For example, this invention uses a dual-path multi-scale interaction method to further improve the model's performance.

[0206] 3. Iteration of conventional semantic information

[0207] In designing the semantic path, this invention uses a semantic tokenization module to apply spatial attention weights to the features obtained from the backbone feature extraction network. Therefore, this invention can achieve technological iteration by updating the specific implementation method in the spatial attention module.

[0208] 4. The present invention provides a dual-path remote sensing image change detection method based on pixel and semantic information interaction. The semantic tags learned by the model from the semantic path further serve as high-level semantics to promote more refined local feature extraction in the pixel path.

[0209] 5. The change detection network constructed in this invention can be divided into three parts. First, a high-resolution pixel path encoder CSCB is used to extract pixel information. Then, high-level semantic information is semantically tokenized to obtain tokens, which are then fed into the Transformer encoder. In the middle, the features after passing through the semantic encoder and pixel encoder respectively enter the pixel and semantic interaction module PSIM. After repeating N times, the dual-temporal pixel features are output. After subtraction, the decoder restores the image to the same size as the input image, which is the prediction result image.

[0210] 6. Unlike semantic segmentation, change detection aims to capture differences between images from different time phases. Therefore, change detection focuses more on information differences between different channels. This characteristic makes the design of change detection methods different from that of semantic segmentation. Current methods, influenced by semantic segmentation, tend to focus on the spatial dimension, using CNNs or Transformers for feature extraction, while paying less attention to channel-dimensional feature extraction. The dual-temporal channel information interaction proposed in this invention provides innovative space for future change detection methods. Therefore, it is expected that by taking a deeper approach from the channel dimension, the perception and analysis capabilities of change detection algorithms for changes in images from different time phases can be improved, bringing new breakthroughs to the field.

[0211] 7. This invention designs a lightweight change detection network that enables rapid training and deployment, improving the accuracy of water change detection with less computing resources.

[0212] The innovative aspects of this invention include:

[0213] 1) A dual-path feature extraction structure was designed, namely a low-level pixel path and a high-level semantic path with multiple branches.

[0214] 2) Design a dual-time-phase interleaved channel information interaction module;

[0215] 3) A pixel semantic information fusion module was designed;

[0216] 4) Design a lighter network structure.

[0217] Experimental Analysis

[0218] Figure 7 The image shows a comparison between the results extracted by this invention in several typical scenarios and seven superior change detection algorithms. The first column shows the image of test phase one, the second column shows the image of test phase two, the third column shows the labels of changes in phase one and phase two, the fourth column shows the extraction results of the FC-EF algorithm, the fifth column shows the extraction results of the USSFC algorithm, the sixth column shows the extraction results of the BIT algorithm, the seventh column shows the extraction results of the SNUNet algorithm, the eighth column shows the extraction results of the OmbriaNet algorithm, the ninth column shows the extraction results of the DESSN algorithm, the tenth column shows the extraction results of the ChangeFormer algorithm, and the eleventh column shows the extraction results of the model proposed in this invention in these scenarios. Green pixels represent falsely detected water bodies, and red pixels represent falsely detected water bodies. The comparison results show that the water extraction model proposed in this invention not only outperforms other methods in large water bodies, small water bodies, reservoirs, and tiny water bodies, but also has very few misclassified water bodies, thus exhibiting the best performance.

[0219] This invention also provides a dual-path remote sensing image change detection system based on pixel and semantic information interaction, comprising:

[0220] Dataset building module: used to obtain the remote sensing image change detection dataset in step 1, and to perform data annotation and image block processing on it to obtain the CDDS dataset;

[0221] Dataset processing module: used to preprocess the CDDS dataset obtained in step 1 in step 2, divide the preprocessed CDDS dataset into training set, validation set and test set, and perform data augmentation on training set and validation set;

[0222] Dual-path change detection network building module: used to implement the construction of the dual-path change detection network PST-Net based on pixel and semantic information interaction in step 3;

[0223] Multiple hyperparameter scheme design module: used to implement the design of multiple different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction constructed in step 3 in step 4;

[0224] The module for obtaining the optimal model under multiple different hyperparameter schemes is used to implement the multiple different hyperparameter schemes designed in step 4 in step 5. The training set and validation set processed in step 2 are used to train the dual path change detection network PST-Net based on pixel and semantic information interaction constructed in step 3 to obtain the optimal model under multiple different hyperparameter schemes.

[0225] The module for obtaining the optimal model under the optimal hyperparameter scheme is used to train the multiple sets of optimal models under different hyperparameter schemes obtained in step 5 using the training set and validation set processed in step 2 in step 6, so as to obtain the optimal model under the optimal hyperparameter scheme.

[0226] Multiple evaluation index acquisition module: This module is used in step 7 to implement the optimal model obtained in step 6, and to test it using the test set divided in step 2 to obtain multiple evaluation indexes.

[0227] The present invention also provides a dual-path remote sensing image change detection device based on pixel and semantic information interaction, comprising:

[0228] Memory: A computer program that stores the above-mentioned dual-path remote sensing image change detection method based on pixel and semantic information interaction, and is a computer-readable device;

[0229] Processor: Used to implement the dual-path remote sensing image change detection method based on pixel and semantic information interaction when executing the computer program.

Claims

1. A dual-path remote sensing image change detection method based on pixel and semantic information interaction, characterized in that: Includes the following steps: Step 1: Obtain the remote sensing image change detection dataset, and perform data annotation and image patching to obtain the CDDS dataset; Step 2: Preprocess the CDDS dataset obtained in Step 1, divide the preprocessed CDDS dataset into training set, validation set and test set, and perform data augmentation on the training set and validation set; Step 3: Construct the PST-Net dual-path change detection network based on pixel and semantic information interaction; The specific method for step 3 is as follows: The dual-path change detection network PST-Net based on pixel and semantic information interaction includes: a backbone feature extraction network, a pixel-semantic dual-path extraction and interaction module, and a change detection head; Step 3.1: Construct the backbone feature extraction network: The backbone feature extraction network is ResNet18. First, the features of phase 1 and phase 2 are passed through the ResNet18 network. After two downsamplings, the features become 1 / 4 of the original. After two more downsamplings, the features become 1 / 16 of the original. The 1 / 4 and 1 / 16 features are then input into the pixel semantic dual-path extraction interaction module for further processing. Step 3.2: Construct a pixel-semantic dual-path extraction interaction module: The pixel semantic dual-path extraction interaction module consists of three parts: a channel scrambling convolution module CSCB, a Transformer encoder, and a pixel semantic interaction module PSIM; Step 3.2.1: Construct the CSCB channel shuffling convolution module: The CSCB channel shuffling convolution module consists of dual-temporal pixel-level feature concatenation, pooling, channel convolution module CCM, and upsampling. The CCM channel convolution module performs channel convolution in the C×HW dimension; the specific operation is as follows: 1) B×C×H×W=>B×C×HW, first straighten the feature map in the spatial dimension; 2) B×C×HW=>B×NC×C×HW, adding a channel dimension, NC, to the second dimension. NC refers to the new channel. After the transformation, a matrix that conforms to the convolution operation is constructed. The initial value of NC is 1. 3) Shuffle: 3×3 convolution is used for sliding to interact information between different channels. Before performing convolution, a channel shuffling operation is performed. 4) After performing convolution, perform inverse shuffle and B×1×C×HW=>B×C×H×W; Step 3.2.2: Construct a multi-layer Transformer encoder: The Transformer encoder consists of block embedding, layer normalization, multi-head attention, and multilayer perceptron. The specific process is as follows: First, the image is embedded into blocks to obtain an image block sequence, where each image block is represented as a vector; then, the image block sequence is input to the normalization layer, followed by multi-head attention calculation and residual connection; finally, after layer normalization and multi-layer perceptron, the output is generated; the Transformer encoder will be stacked multiple times to form a multi-layer Transformer encoder. Step 3.2.3: Construct the PSIM pixel semantic interaction module: The PSIM pixel semantic interaction module overlaps N layers of multi-head attention. The structure includes 4 LayerNorms, 2 feedforward networks (FFN), and 2 multi-head self-attention. The module has two inputs, which are initially features from the pixel path and the semantic path, respectively. The pixel path is used to extract low-level features of the image, while the semantic path focuses on high-level information in the image. Step 3.3: Construct the change detection head; At this point, the pixel features of phase 1 and phase 2, after N layers of interaction in the PSIM pixel semantic interaction module, contain rich semantic information. After subtraction, the difference map of the two phases is obtained. The change detection head then performs two 3×3 convolutions on the obtained difference map and directly upsamples it by 4 times to become a change map with the same size as the original input, which is the prediction result map. Step 4: Design multiple sets of different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction constructed in Step 3; Step 5: Based on the multiple sets of different hyperparameter schemes designed in Step 4, use the training set and validation set processed in Step 2 to train the dual-path change detection network PST-Net built in Step 3 based on pixel and semantic information interaction, and obtain the optimal model under multiple sets of different hyperparameter schemes. Step 6: Use the training set and validation set processed in Step 2 to train the optimal models under multiple different hyperparameter schemes obtained in Step 5, and obtain the optimal model under the optimal hyperparameter scheme. Step 7: Using the optimal model obtained in Step 6, test it using the test set divided in Step 2 to obtain various evaluation metrics.

2. The dual-path remote sensing image change detection method based on pixel and semantic information interaction according to claim 1, characterized in that: The specific method for step 1 is as follows: Step 1.1: Collect remote sensing images of the same area at two different time points; Step 1.2: Data Labeling: Use the Labelme tool to annotate the remote sensing images obtained in step 1.1, mark the areas that have changed, and create labels; Step 1.3: Image segmentation: The remote sensing images obtained in step 1.1 and the labels obtained in step 1.2 are divided into blocks. The blocks of images and the corresponding labels are combined to form a dataset, namely the CDDS dataset.

3. A dual-path remote sensing image change detection method based on pixel and semantic information interaction according to claim 1 or 2, characterized in that: The specific method for step 2 is as follows: Step 2.1: Normalize the CDDS dataset obtained in Step 1.3: Calculate the overall maximum and minimum values ​​for the same channel across all images, following the channel dimension. Then, normalize the data for that channel using these overall maximum and minimum values, using the following formula: Formula 1 Formula 2 in, This represents the value of the j-th pixel in the c-th channel of the i-th image. express Normalized pixel values, The minimum pixel value of the c-th channel layer. The maximum pixel value of the c-th channel layer; Step 2.2: After normalization, layers with the same channel in all images are grouped together. Within each group, the population mean and population variance of the data obtained in Step 2.1 are calculated. The formulas for calculating the population mean and population variance are as follows: Formula 3 Formula 4 in, This represents the pixel mean of the c-th channel. This represents the standard deviation of pixels in the c-th channel. Let N represent the value of the j-th pixel in the c-th channel of the i-th image, N represent the number of images, and M represent the number of pixels in a single image. Step 2.3: Data standardization processing: For all the normalized data obtained in step 2.2 Standardization is performed at the channel level, using the following formula: Formula 5 in, This represents the standardized value of the j-th pixel in the c-th channel of the i-th image; Step 2.4: Divide the standardized data from Step 2.3 into training, validation, and test sets; Step 2.5: Perform data augmentation on the training and validation sets from Step 2.4: We use "vertical flip + horizontal flip", "shift", and "scaling" to augment the training and validation sets.

4. The dual-path remote sensing image change detection method based on pixel and semantic information interaction according to claim 1, characterized in that: The specific method for step 4 is as follows: Step 4.1: For the four hyperparameters Layer, STL, Dim and KS in the dual-path change detection network PST-Net based on pixel and semantic information interaction constructed in Step 3, design multiple sets of different values ​​for each. Step 4.2: Using the control variable method, multiple sets of values ​​for the four hyperparameters in Step 4.1 are combined to obtain multiple different hyperparameter schemes for the dual-path change detection network PST-Net based on pixel and semantic information interaction.

5. A dual-path remote sensing image change detection method based on pixel and semantic information interaction according to claim 1 or 2, characterized in that: The specific method for step 5 is as follows: Using the training and validation sets in the CDDS dataset processed in step 2, we trained and tested multiple hyperparameter schemes of the dual-path change detection network PST-Net obtained in step 4.2 based on pixel and semantic information interaction. We analyzed the test results and finally obtained the optimal hyperparameter scheme through model evaluation.

6. The dual-path remote sensing image change detection method based on pixel and semantic information interaction according to claim 1, characterized in that: The specific method for step 6 is as follows: Step 6.1: Construct the dual-path change detection network PST-Net based on pixel and semantic information interaction in Step 3 according to the optimal hyperparameter scheme obtained in Step 5; Step 6.2 Train the PST-Net network constructed in Step 6.1 using the training and validation sets processed in Step 2; during training, the cross-entropy (BCE) loss function is used. Defined as: Formula 6 in, It's a real label. The predicted probability of the model; Step 6.3: During each training process, save the model with the highest F1 coefficient on the validation set after Step 2 as the optimal model.

7. The dual-path remote sensing image change detection method based on pixel and semantic information interaction according to claim 1, characterized in that: The specific method for step 7 is as follows: Based on the optimal model with the highest F1 coefficient obtained in step 6, test it using the test set from step 2. Use the model's F1 score, pixel accuracy PA, mean cross-union ratio mIoU, frequency weight cross-union ratio FWIoU, and Kappa coefficient as evaluation indicators. The calculation formulas for the evaluation indicators F1, pixel accuracy PA, mean intersection-over-union ratio mIoU, weighted intersection-over-union ratio FWIoU, and Kappa coefficient are as follows: Formula 7 Formula 8 Formula 9 Formula 10 Formula 11 Formula 12 The meanings of the parameters in equations 7, 8, 9, 10, 11, and 12 are as follows: TP: Prediction correct, predicted result is positive, and the actual value is positive. FP: Prediction incorrect, predicted result is positive, and the actual value is negative. FN: Prediction incorrect, predicted result is negative, and the actual value is positive. TN: Prediction correct, predicted result is negative, and the actual value is negative.

8. A dual-path remote sensing image change detection system based on pixel and semantic information interaction, used to implement the method of claim 1, characterized in that: include: Dataset building module: used to acquire remote sensing image change detection dataset, and to perform data annotation and image patching processing to obtain CDDS dataset; Dataset processing module: Used to preprocess the CDDS dataset, divide the preprocessed CDDS dataset into training set, validation set and test set, and perform data augmentation on the training set and validation set; Dual-path change detection network building module: used to build the dual-path change detection network PST-Net based on pixel and semantic information interaction; Multiple hyperparameter scheme design module: used to design multiple different hyperparameter schemes for the constructed dual-path change detection network PST-Net based on pixel and semantic information interaction; The module for obtaining the optimal model under multiple different hyperparameter schemes is used to train the constructed dual-path change detection network PST-Net based on pixel and semantic information interaction using the processed training and validation sets according to multiple different hyperparameter schemes, so as to obtain the optimal model under multiple different hyperparameter schemes. The module for obtaining the optimal model under the optimal hyperparameter scheme is used to train multiple sets of optimal models under different hyperparameter schemes using the processed training set and validation set, and obtain the optimal model under the optimal hyperparameter scheme. Multiple evaluation metrics acquisition module: Used to obtain multiple evaluation metrics by using the optimal model and the partitioned test set for testing.

9. A dual-path remote sensing image change detection device based on pixel and semantic information interaction, characterized in that: include: Memory: A computer program for a dual-path remote sensing image change detection method based on pixel and semantic information interaction as described in any one of claims 1-7, and is a computer-readable device; Processor: Used to implement the dual-path remote sensing image change detection method based on pixel and semantic information interaction as described in any one of claims 1-7 when executing the computer program.

Citation Information

Patent Citations

  • Two-stage high-resolution remote sensing image change detection method in technical field of remote sensing

    CN110263705A

  • Transform and dense feature fusion-based remote sensing image change detection method and system

    CN115690002A