Change Detection Method Based on the Hybrid of Convolutional Network and Graph Neural Network
By combining convolutional networks and graph neural networks in the change detection algorithm, local and global features of remote sensing images are extracted, and feature fusion is used to use point-to-point feature fusion modules to perform feature fusion, the problems of high computational complexity and high time overhead when extracting global features in the prior art are solved, and more efficient feature fusion and change detection effects are achieved.
Patent Information
- Application Number
- CN202310570042.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-05-19
AI Technical Summary
The existing change detection algorithm has problems with high computational complexity and high time overhead when extracting global features of remote sensing images, which leads to the inability to fully utilize the potential of Transformer to extract global features.
Using a change detection method based on the hybrid of convolutional network and graph neural network, multi-scale semantic features are extracted through the encoder, the fusion device performs feature fusion, the decoder gradually recovers the change information, and uses the point-to-point feature fusion module (P2PFFM) to perform feature fusion to reduce time overhead.
It realizes the simultaneous extraction of local and global features, reducing time overhead, making network optimization more clear and focused, thereby better completing feature fusion and improving the effect of change detection.
Smart Images

Figure CN116778317B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of remote sensing image change detection, specifically a change detection method based on a hybrid of convolutional network and graph neural network. Background Art
[0002] In recent years, almost all computer vision algorithms are based on convolutional neural networks, and change detection algorithms are no exception.
[0003] Most of the previous change detection algorithms are based on convolutional neural networks. In recent years, many algorithms are also based on hybrid networks of convolution and Transformer. The change detection network based on convolution is restricted by the locality of convolution and cannot effectively extract global features in space. For the hybrid change detection network based on convolution and Transformer, although it utilizes Transformer, due to the quadratic relationship between the computational complexity of Transformer and the input size and its high time overhead, Transformer is usually only used as a module in the network and cannot fully exert the potential of Transformer to extract global features. Summary of the Invention
[0004] To solve the deficiencies of the prior art, the present invention provides a change detection method based on a hybrid of convolutional network and graph neural network, which can extract local features and global features simultaneously, has a lower time overhead, makes the optimization of the network more clearly focused, and thus better completes feature fusion.
[0005] To achieve the above object, the present invention is realized through the following technical solutions:
[0006] A change detection method based on a hybrid of convolutional network and graph neural network includes an encoder, a fuser, and a decoder.
[0007] The encoder has a pyramid structure with a total of five layers. As the number of layers deepens, the spatial size of the feature map decreases, and the spatial size of each layer is reduced by half compared to the previous layer. At the same time, the number of channels will also increase accordingly. The input of the encoder is the remote sensing image to be processed, denoted as X Ein ∈R H×W×3 After being processed by these five layers, the feature extracted by the encoder is (C Eout is the number of channels);
[0008] The fuser receives the output feature X Eout1 of the encoder and X Eout2, they are merged. The features input to the fuser first pass through a concatenation layer, where the input dual features are concatenated in the channel dimension. Then, a linear mapping layer reduces the number of channels of the features. The linear mapping layer consists of a fully connected layer. Finally, the fuser uses L F visual graph convolutional blocks to further better merge information. In the fuser, the sizes of the input and output features are both
[0009] The deep network of the decoder has a total of 3 layers. Each layer consists of an upsampling layer, a point-to-point feature fusion module, and 2 visual graph convolutional modules. The shallow network of the decoder has 2 layers. Each layer mainly consists of an upsampling layer, a point-to-point feature fusion module, and a convolutional module. The last layer also includes a linear mapping layer for classification;
[0010] Above, H and W are the height and width of the image respectively, and C Eout is the number of channels.
[0011] Specifically, it includes the following steps:
[0012] S1), The encoder extracts effective multi-scale semantic features from the input dual-time images respectively;
[0013] S2), The fuser effectively fuses the dual-time features generated by the encoder together for further processing by the decoder;
[0014] S3), The decoder gradually recovers the change information from the extracted features and finally generates a change image.
[0015] In the above detection task, there is a skip connection that introduces the features generated by the encoder into the decoder. In the decoder, the introduced features and the features in the decoder are fused through a point-to-point feature fusion module. For multiple feature maps, the point-to-point feature fusion module uses the points in the same position to update each point, so that the points of multiple feature maps correspond to each other, forming a better complementary relationship.
[0016] Compared with the prior art, the beneficial effects of the present invention are:
[0017] 1. The present invention proposes a new change detection network framework, a change detection network that combines convolution and graph convolution, namely HCGNet. HCGNet uses convolution operations in the shallow network and graph convolution operations in the deep network. The features extracted by the shallow network are mainly some detailed features. The convolutional network has inherent locality and can well extract local features with detailed information. The deep network mainly extracts the semantic features of the network and needs to extract information in the long-term space. The graph convolutional network finds nodes with similar content for all nodes globally and uses the similar nodes to update each node, so as to be able to extract semantic features globally and with less time overhead. HCGNet cleverly utilizes the respective characteristics of convolution and graph convolution to construct an effective change detection network.
[0018] 2. The present invention proposes a new feature fusion module, a point-to-point feature fusion module, namely P2PFFM. HCGNet uses P2PFFM to fuse the features of the decoder and the encoder. P2PFFM calculates the weights of each point on the feature map from three dimensions: channel, width, and height, and uses the points at the same position in all feature maps to update each point. P2PFFM makes the points at the same position in all feature maps easier to form corresponding and complementary relationships, making the optimization of the network clearer and more focused.
[0019] In summary, the present invention adopts a hybrid network, using a convolutional network to extract local features in the shallow layer and a graph neural network to extract global features in the deep layer. The technical effect is obvious, and it can extract local features and global features, thus better completing feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG Figure 1 is a schematic diagram of the change detection network of the present invention;
[0021] FIG Figure 2 is a schematic diagram of the encoder structure of the present invention;
[0022] FIG Figure 3 is a schematic diagram of the fusion structure of the present invention;
[0023] FIG Figure 4 is a schematic diagram of the decoder structure of the present invention;
[0024] FIG Figure 5 is a schematic diagram of the P2PFFM feature fusion method of the present invention;
[0025] FIG Figure 6 is a network structure diagram of the P2PFFM of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0027] As Figure 1 shown, the change detection network based on the hybrid of convolutional neural network and graph neural network, namely HCGNet, in order to be able to process the dual-temporal change images simultaneously, HCGNet has a twin feature extraction process, and at the same time HCGNet is also a typical U-shaped structure.
[0028] Overall, HCGNet is divided into three parts, named respectively: encoder, fuser and decoder, and skip connections are constructed between the encoder and the decoder, and the skip connections will mix the features generated by the encoder with the features in the decoder. When the skip connections connect the features of the encoder and the decoder, HCGNet proposes a point-to-point feature fusion module, namely P2PFFM. For each point of multiple feature maps, P2PFFM uses the points in the same position to update it. The features are fused in the dimension of the points, the operation is more concise, and the goal is more specific.
[0029] Regarding the network structure, HCGNet uses convolutional blocks as the basic building units in the shallow layer, and uses visual graph convolutional blocks as the basic building units in the deep layer. In the shallow layer of the network, the feature maps usually have a large size, and lightweight operation units are required. Operations that can extract global features, such as Transformer, usually have huge time overhead and are not suitable for extracting features of larger sizes. At the same time, the shallow feature maps contain a large amount of detailed information, such as texture, edge and other information, which is suitable for extracting local features. The convolutional operation has inherent locality. Therefore, HCGNet uses a lightweight convolutional operation to construct the shallow network. The graph convolutional module regards the feature maps as nodes one by one, constructs a graph relationship according to the similarity relationship of the nodes, and updates each node using the information of the connected nodes. The graph convolutional module searches for nodes with similar content for each node globally. Therefore, the graph convolutional module is not restricted by space and can construct long-term features. The deep network is usually suitable for extracting the semantic features of images in the long-term space. Therefore, the deep network of HCGNet is mainly composed of graph convolutional modules. The shallow convolutional structure and the graph convolutional structure in HCGNet complement each other and can better extract the local and global features of images.
[0030] Specifically, regarding the encoder, the function of the encoder in the present invention is to extract effective multi-scale semantic features from the input dual-moment images. The encoder of HCGNet uses a convolutional module to extract features in the shallow network and a graph neural network module to extract features in the deep network. The reason for doing this is that: the features extracted by the shallow network have a lot of detailed information, and the convolutional module has local characteristics, which can better capture the detailed information in the shallow network. The deep network is more suitable for extracting the global features of the image, and the graph neural network module can establish associations between features with similar content, regardless of distance, so as to better and more efficiently extract semantic features with global information in the deep network.
[0031] The encoder has a pyramid structure and it has a total of five layers. As the number of layers deepens, the spatial size of the feature map also decreases. The spatial size of each layer is reduced by half compared to the previous layer, and at the same time the number of channels will also increase accordingly. The input of the encoder is the remote sensing image to be processed, denoted as X Ein ∈R H×W×3 (where H and W are the height and width of the image respectively). After being processed by these five layers, the features extracted by the encoder are (C Eout is the number of channels). The encoder processes the dual-moment images simultaneously, so two output features are obtained, denoted as X Eout1 and X Eout2 respectively.
[0032] The first two layers of the encoder belong to the shallow network of the encoder, and the specific structure is as Figure 2 shown in Fig. a. In the present invention, the shallow network of the encoder consists of two branches, a convolutional branch and a patch embedding branch. The convolutional branch includes two layers, which are mainly composed of convolutional blocks and downsampling layers. The shallow network of the encoder of the ViG
[26] network consists of a patch embedding module, which performs downsampling on the image and is composed of convolutional layers. To be consistent with the ViG network, the shallow network of the encoder also includes the same patch embedding branch. The results of the convolutional branch and the patch embedding branch will go through a summation operation to jointly obtain the output of the shallow network. The latter three layers belong to the deep network of the encoder, and the structure is as Figure 2 shown in Fig. b. Each layer consists of L Ei visual graph convolutional blocks and a downsampling layer, where i represents the index of the layer.
[0033] Regarding the fuser, as mentioned above, in the encoder, the input dual-moment images are processed simultaneously, and the encoder generates dual-moment features. The function of the fuser is to effectively fuse the dual-moment features generated by the encoder for further processing by the decoder. The fuser receives the output features X Eout1 and XEout2 , they are merged. The structure of the fuser is as Figure 4 shown. The input features first pass through a concatenation layer, and the input dual features are concatenated in the channel dimension. Then, a linear mapping layer reduces the number of channels of the features. The linear mapping layer consists of a fully connected layer. Finally, the fuser uses L F visual graph convolutional blocks to further merge information better. In the fuser, the sizes of the input and output features are both
[0034] Regarding the decoder, the function of the decoder in the present invention is to gradually recover the change information from the extracted features and finally generate a change image. The decoder has a structure symmetric to that of the encoder, an inverted pyramid structure, as Figure 4 shown. The deep network of the decoder has a total of three layers, and each layer consists of an upsampling layer, a P2PFFM, and two visual graph convolutional modules. The shallow network of the decoder has two layers, and each layer mainly consists of an upsampling layer, a P2PFFM, and a convolutional module. The last layer also includes a linear mapping layer for classification, that is, a fully connected layer. Each time the feature map passes through an upsampling layer, its size doubles and the number of channels decreases accordingly. After these five levels of processing, the size of the feature map changes from to H×W×2. The output X Dout of the decoder is the predicted change map.
[0035] Regarding P2PFFM, specifically, in the change detection task, there is usually a skip connection to introduce the features generated by the encoder into the decoder. In the decoder, the introduced features need to be fused with the features in the decoder. To better complete this fusion operation, the present invention proposes a new feature fusion module, the point-to-point feature fusion module, that is, the above-mentioned P2PFFM. For multiple feature maps, P2PFFM uses the points at the same position to update each point, as Figure 5 shown. In the figure, the green points and the star points have the same position, and P2PFFM uses the values of these green points to update the values of the star points. P2PFFM makes the points of multiple feature maps correspond to each other, forming a better complementary relationship.
[0036] The structure of P2PFFM is as Figure 6 shown. In HCGNet, the input of P2PFFM is three feature maps, where represents the features in the decoder, and Represents the features generated by the corresponding layer of the encoder. i and j ∈ [1, 5] respectively represent the indices of the decoder and encoder layers where the features are located, and j = 5 - i + 1. Their size is denoted as h × w × c, where h, w, and c represent the height, width, and number of channels of the feature map respectively. P2PFFM performs point-to-point feature fusion on these three feature maps in the three dimensions of channels, height, and width respectively. Taking the first stage of P2PFFM, that is, the point-to-point feature fusion in the channel dimension, as an example to explain the fusion process. First and will be concatenated in the channel dimension, and the size of the concatenated feature is h × w × 3c. The subsequent operations are divided into two branches. The first branch includes a multi-layer perceptron operation, a matrix dimension transformation operation, and a softmax operation. The multi-layer perceptron operation reduces the number of channels of the feature from 3c to 9. The matrix dimension transformation operation adjusts the size of the feature from h × w × 9 to h × w × 3 × 3. The softmax operation adjusts the values of the feature to between 0 and 1. The second branch contains a matrix dimension transformation operation that adjusts the size of the feature to h × w × 3 × c. The results of the two branches will be subjected to a matrix multiplication operation, thus completing the feature fusion operation in the channel dimension. After that, P2PFFM has two more feature fusion operations, which are to fuse the features in the height and width dimensions respectively, and the specific operations are similar to the feature fusion in the channel dimension. After the feature fusion is completed, a matrix dimension transformation operation will be used to change the size of the feature to h × w × 3c. Finally, a linear mapping layer (fully connected layer) reduces the number of channels to c.
[0037] Example:
[0038] To verify the effect of the present invention, in this embodiment, the CDD dataset contains 11 pairs of bi-temporal images with seasonal variations, including 7 pairs of images with 4725×2200 pixels and 4 pairs of images with 1900×1000 pixels. These images were taken by Google Earth with a spatial resolution of 3 - 100 cm / pixel. In the CDD dataset, the changes to be detected include changes caused by buildings, cars, roads, etc., but do not include changes caused by seasons. All images were segmented into 256×256 image patches through cropping and rotation, generating 16000 pairs of image patches. Among these image patches, 10000 pairs were used for training, 3000 pairs were used for validation, and the remaining were used for testing.
[0039] LEVIR-CD is also a dataset collected through Google Earth. It was taken in Texas, USA from 2002 to 2018, focusing on building changes. LEVIR-CD contains 637 groups of images with a size of 1024×1024 pixels. Each group contains a pair of bi-temporal images and a label image. 64 groups are used for validation, 445 groups for training, and 128 groups for testing. The spatial resolution of these images is 0.5m / pixel.
[0040] In this embodiment, four evaluation metrics are used to evaluate the test results. They are precision, recall, F1-score, and overall accuracy (OA), where the F1-score is the main evaluation metric. The calculation formulas are as follows:
[0041] Precision = TP / (TP + FP)
[0042] Recall = TP / (TP + FN)
[0043]
[0044]
[0045] TP is the number of true positive samples, TN is the number of true negative samples, FP is the number of false positive samples, and FN is the number of false negative samples.
[0046] As shown in the following table, by conducting experiments on two datasets, LEVIR and CDD, the F1 value of HCGNet is at least 0.74% and 1.1% higher than other change detection methods respectively.
[0047] Method LEVIR CDD FC-EF 80.60 57.45 CDNet 84.17 85.95 STANet 84.87 93.48 DASNet 85.18 92.53 SNUNet 90.28 94.65 BiT 89.78 95.28 ChangeFormer 89.59 94.72 DARNet 90.42 96.31 HCGNet 91.16 97.41
Claims
1. A change detection method based on a hybrid of a convolutional network and a graph neural network, including an encoder, a fuser, and a decoder, characterized in that: The encoder has a pyramid structure with a total of five layers. As the layer depth increases, the spatial size of the feature map decreases, with the spatial size of each layer being half that of the previous layer. At the same time, the number of channels also increases accordingly. The input to the encoder is the remote sensing image to be processed, denoted as X Ein ∈R H×W×3 After being processed through these five layers, the features extracted by the encoder are The fuser receives the output feature X of the encoder Eout1 and X Eout2 , combines them, and the features input to the fuser first pass through a concatenation layer, where the input dual features are concatenated in the channel dimension. Then, a linear mapping layer reduces the number of channels of the features. The linear mapping layer consists of a fully connected layer. Finally, the fuser uses L F visual graph convolutional blocks to further better combine information. In the fuser, the sizes of the input and output features are both The deep network of the decoder has a total of 3 layers, and each layer consists of an upsampling layer, a point-to-point feature fusion module, and 2 visual graph convolutional modules. The shallow network of the decoder has 2 layers, and each layer mainly consists of an upsampling layer, a point-to-point feature fusion module, and a convolutional module. The last layer also includes a linear mapping layer for classification; Among the above, H and W are the height and width of the image respectively, and C Eout is the number of channels; Specifically, it includes the following steps: S1), The encoder extracts effective multi-scale semantic features from the input dual-moment images respectively; S2), The fuser effectively fuses the dual-moment features generated by the encoder together for further processing by the decoder; S3), The decoder gradually recovers the change information from the extracted features and finally generates a change image; In the above detection task, there is a skip connection that introduces the features generated by the encoder into the decoder. In the decoder, the introduced features and the features in the decoder are fused through a point-to-point feature fusion module. For multiple feature maps, the point-to-point feature fusion module uses the points in the same position to update each point, so that the points of multiple feature maps correspond to each other, forming a better complementary relationship.