SAR Image Change Detection Method Based on Hybrid Convolution
Multi-scale features are extracted through graph convolution enhancement convolution network (GECN) and combined with graph volume neural network and progressive fusion module, the shortcomings of extracting multi-scale features and utilizing global information in the existing technology are solved, and high-precision SAR image change detection is achieved and computing resource consumption is reduced.
Patent Information
- Application Number
- CN202211472616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The existing SAR image change detection technology has shortcomings in extracting multi-scale features and utilizing global information, resulting in low detection accuracy and excessive computing resource consumption.
Graph convolution enhancement convolution network (GECN) is used to extract multi-scale features, obtain global information through graph volume neural network submodules, and combine the progressive fusion module and label update module to achieve effective fusion of features and label optimization.
It improves the accuracy of SAR image change detection, reduces the consumption of computing resources, and can more effectively explore the multi-scale and global information of the image.
Smart Images

Figure CN115761502B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing and relates to a method for detecting changes in synthetic aperture radar (SAR) images, which can be used for urban planning and layout, natural disaster assessment and military dynamic reconnaissance. Background Art
[0002] Change detection in SAR images plays a vital role in the field of remote sensing and has attracted increasing attention due to its wide range of applications. Existing SAR image change detection techniques can be divided into two categories.
[0003] The first category is the change detection method based on the traditional algorithm. In the region recognition stage, this method uses the feature difference and ratio of pixel pairs or object pairs as input and detects changes through thresholds. According to the size of the image processing unit, it can be further divided into object-based methods and pixel-based methods. However, whether it is pixel-based or object-based change detection method, only manually extracted features such as spectral features, texture features, and shape features are used to extract change information. Since these features cannot fully represent the key information of SAR images, they greatly affect the change detection results.
[0004] The second category is change detection based on deep learning algorithms. This method uses the powerful nonlinear fitting ability and multi-level structure of convolutional neural networks to make the learned features have high-level semantic information and rich spatial context information. Therefore, it is a widely used change detection method. Although the change detection method based on deep learning has achieved great success in the change detection task of SAR images, there are still some problems to be further solved. First, only multi-level convolutional neural networks are used to extract the features of SAR images. In theory, deep features have global context information, but in fact, the receptive field of deep features is not global. At the same time, global information cannot be effectively used on shallow features. Therefore, there is a lack of global information to model the features of the image. Secondly, the fusion of SAR multi-scale features is generally carried out by upsampling and downsampling to the same size, and then directly adding or splicing channels. However, in practice, multi-scale features cannot be well fused after only one processing, which will make the features biased towards a certain scale feature and fail to make full use of multi-scale features. Finally, most unsupervised change detection algorithms use clustering to obtain pseudo labels for training, so the selected training labels are mixed with a large number of wrong labels, and their credibility is low, so that network training cannot achieve the expected effect.
[0005] In the paper on heterogeneous remote sensing change detection based on nonlocal patch similarity published by Sun Yuli in Science Citation Index (SCI) 2021, 109: 107598, the nonlocal patch similarity based graph (NPSG) was proposed to measure the structural consistency between heterogeneous images, thereby completing the change detection task. The specific implementation is as follows: First, the image is divided into a series of patches. In the forward detection, for each target patch in the pre-event image, its k-nearest NPSG in the pre-event image is calculated using a statistical similarity metric, and then this is mapped to the nearest NPSG of the post-event image, and its own nearest NPSG of the post-event image is compared with the nearest NPSG of the pre-event image by calculating the similarity difference; then, the change detection result is obtained by using the threshold segmentation method for the comparison result. This method uses the k-nearest neighbor graph structure to model the global contextual relationship of the target patch. Although it can achieve good change detection accuracy, it requires a lot of computing resources when calculating the similarity between different patches.
[0006] Dong Huihui et al. proposed a SAR image change detection method based on multi-scale self-attention deep clustering MSDC in the Institute of Electrical and Electronics Engineers IEEE, 2021, 60: 1-16. First, the side window filtering algorithm is used to reduce the speckle noise in the SAR image; then, the multi-scale adjacent areas of the given central pixel of the processed image are used as the basic difference analysis unit of the sample; then, the multi-scale features are fused to obtain the difference features of the two branches, and finally, the network is optimized by clustering to obtain the labels. In the process of multi-scale feature fusion, the convolution nonlinear transformation is first used to adjust the features of the three scales to the same size, and then the three features of the same size are regarded as the query q, key value k, and value v in the self-attention module. The self-attention mechanism is used to obtain a feature; finally, the feature obtained by self-attention is added to the feature before self-attention to complete the fusion of multi-scale features. Although Dong Huihui's ablation experiment proves the effectiveness of multi-scale attention, the generation of the weight matrix in the self-attention mechanism requires a lot of computing resources to support, resulting in too high computing cost and inconvenient application.
[0007] Qu Xiaofan et al. published a change detection method for SAR images using a dual-domain network DDNet in IEEE, 2021, 19: 1-5. It uses feature representations in the frequency and spatial domains to mitigate speckle noise, uses hierarchical fuzzy clustering to obtain labels, and selects 10% of reliable labels based on membership to train the network. Although this method can effectively reduce the number of incorrect labels, the small number of labels involved in training hinders further improvement of network performance. Summary of the invention
[0008] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a SAR image change detection method based on hybrid convolution to extract multi-scale features with global and local information, improve detection accuracy and reduce computing resources.
[0009] To achieve the above-mentioned purpose, the technical solution of the present invention comprises the following steps:
[0010] (1) Build a graph convolution enhanced convolutional network GECN:
[0011] 1a) Establishing a feature extraction module consisting of a convolutional neural network submodule CNNB, a graph convolution neural network submodule GCNB and a fusion submodule, wherein CNNB and GCNB are connected in parallel and then in series with the fusion submodule;
[0012] 1b) Establishing a progressive fusion module that sequentially performs up- and down-sampling operations, concatenation operations, and convolution operations;
[0013] 1c) Establishing a label update module that updates the training labels by voting with the network prediction labels and prototype labels;
[0014] 1d) connecting the feature extraction module, the progressive fusion module and the label update module in series to form a graph convolution enhanced convolutional network GECN;
[0015] (2) Obtain GECN input and initial label:
[0016] 2a) A logarithmic ratio operator is applied to the dual-time SAR image to obtain a difference map, and the difference image is concatenated with the dual-time image in the channel dimension to obtain the input of the three-channel GECN;
[0017] 2b) clustering the difference graph in 2a) using a fuzzy clustering algorithm to obtain an initial label;
[0018] (3) Use GECN input and labels to train GECN:
[0019] 3a) In the initial 20 training batches, the three-channel input is fed into GECN, the cross entropy loss function is applied to the GECN output and the initial label, and the GECN parameters are updated using the gradient descent method to complete the GECN pre-training;
[0020] 3b) In the subsequent 50 training batches, the three-channel input is fed into the pre-trained GECN, the initial label is updated using the label update module, the cross entropy loss function is applied to the output of the pre-trained GECN and the updated initial label, and the gradient descent method is used to further update the pre-trained GECN parameters to complete the training of the GECN;
[0021] (4) The test SAR image is input into the trained GECN to obtain the GECN output. The max function is used on the GECN output to obtain the category of each pixel sample on the SAR image, which is the final change detection result of the SAR image.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] 1) The present invention uses the graph convolution neural network submodule GCNB to take the point vectors on the compressed features as graph nodes, and performs graph convolution operations along the x-axis and y-axis directions respectively. Therefore, the global information of the image can be obtained with little information loss. It is not only simple to operate, but also can fully mine the global information of the image.
[0024] 2) Since the present invention adopts an adaptive fusion method to fuse the local and global information of the image, it can more fully explore the information useful for the change detection task.
[0025] 3) Since the present invention embeds the difference map into the input, it is equivalent to providing a priori information for GECN training, which makes GECN converge faster and reduces the occupation of computing resources during training.
[0026] 4) The present invention uses a step-by-step multiple fusion approach to fully explore the multi-scale information of the image and avoid the problem of features being biased towards a certain scale.
[0027] 5) The present invention uses a label update module to update the initial labels, and obtains more and more reliable labels. Using these labels to train GECN can increase the generalization of GECN and improve the change detection accuracy of GECN.
[0028] Simulation experiments show that the present invention can achieve excellent change detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the specific implementation of the present invention or the technical solution in the prior art, the drawings required for use in the specific implementation or the description of the prior art are briefly introduced below.
[0030] Figure 1 It is a flow chart of the implementation of the present invention;
[0031] Figure 2 It is a graph convolution enhanced convolutional network GECN framework diagram in the present invention;
[0032] Figure 3 It is a schematic diagram of the structure of the graph convolutional neural network submodule GCNB of the GECN network in the present invention;
[0033] Figure 4It is a schematic diagram of the fusion submodule structure of the GECN network in the present invention;
[0034] Figure 5 It is a schematic diagram of a label update module in a GECN network of the present invention;
[0035] Figure 6 The figures are simulation results of SAR image change detection using the present invention and the existing detection method respectively. DETAILED DESCRIPTION
[0036] The examples and effects of the present invention are further described in detail below with reference to the accompanying drawings.
[0037] Reference Figure 1 The SAR image change detection method based on hybrid convolution in this example includes four parts: obtaining three-channel input and initial labels, building a graph convolution enhanced convolution network GECN, training GECN using three-channel input and labels, and obtaining the change result map. The specific implementation is as follows:
[0038] Step 1: Get the three-channel input and initial labels.
[0039] 1.1) Construct three-channel input:
[0040] Assume that there are two registered SAR images I acquired at the same location but at different times. 1 and I 2 , in order to identify I 1 and I 2 In order to reduce the influence of speckle noise in SAR images on change detection performance, we first need to 1 and I 2 The logarithmic ratio operator is used to smooth the speckle noise and obtain the difference map D. In order to compensate for the information loss of D caused by the smoothing effect of the logarithmic ratio operator, I 1 ,I 2 The difference map D is concatenated in the channel dimension to obtain a three-channel input.
[0041] 1.2) Get the initial label:
[0042] The difference graph D in 1.1) is clustered using the fuzzy clustering algorithm, and the working process is as follows:
[0043] 1.2.1) Determine the number of fuzzy clustering categories c and the number of iterations t;
[0044] 1.2.2) Initialize the membership matrix u with random numbers between 0 and 1 so that its sum is 1:
[0045]
[0046] Among them, u ij is the membership of the j-th pixel sample to the i-th category, and n is the number of pixel samples;
[0047] 1.2.3) Calculate the cluster center based on the degree of membership:
[0048]
[0049] Among them, c i is the i-th cluster center, x j is the j-th pixel sample;
[0050] 1.2.4) Calculate the sum of squares of errors J from pixel samples to various center points:
[0051]
[0052] 1.2.5) Find the extreme value of formula 1.2.4) and when the sum of squared errors is minimized, we get the updated membership matrix u ij ;
[0053] 1.2.6) Repeat 1.2.3) to 1.2.5) until the iteration number t is reached and the initial label is obtained.
[0054] Step 2: Build the graph convolution enhanced convolutional network GECN.
[0055] Reference Figure 2 ,GECN includes a feature extraction module, a progressive fusion module and a label update module connected in series. The construction process of each module is as follows:
[0056] 2.1) Establish a feature extraction module consisting of a convolutional neural network submodule CNNB, a graph convolutional neural network submodule GCNB and a fusion submodule:
[0057] 2.1.1) Connect three convolutional layers with kernel sizes of 1×1, 3×3, and 1×1 in sequence to form a CNNB submodule to explore the local information of the image;
[0058] 2.1.2) Establish GCNB submodule:
[0059] Reference Figure 3 The GCNB submodule contains two parallel pooling layers and four graph convolution layers. These four graph convolution layers are connected in series with each pooling layer. The working process of the GCNB submodule is as follows:
[0060] First, two parallel pooling layers are used to pool the input F of GCNB along the x-axis and y-axis respectively. in Perform global average pooling to obtain the compressed feature F x and F y :
[0061] F x =Pooling x (F in )
[0062] F y =Pooling y (F in )
[0063] Among them, Pooling x and Pooling y Represents global average pooling along the x-axis and y-axis respectively;
[0064] Next, the compressed feature F x and F y The point vectors on the graph are used as graph nodes to establish F x and F y The adjacency matrix A between the nodes in the above figure x and A y :
[0065]
[0066]
[0067] Among them, γ = 100 is the scaling factor set empirically, dis(·) represents the Euclidean distance between two graph nodes, and v x,i ,v x,j Represent any two x The obtained graph node, v y,i ,v y,j Represent any two y The obtained graph nodes, and Respectively represent F x and F y The graph node space on ;
[0068] Then, according to F x and F y The adjacency matrix A of the nodes in the above figure x The pixel at position (i, j) and F y The adjacency matrix A of the nodes in the above figure y The pixel at position (i, j) Find the corresponding degree matrix and
[0069]
[0070]
[0071] Then, using the degree matrix and For the adjacency matrix A x and A y Perform Laplace normalization to obtain the normalized adjacency matrix L x and L y :
[0072]
[0073]
[0074] Among them, A′ x =A x +I, A' y =A y +I, I is the identity matrix;
[0075] Finally, two layers of graph convolution are used to propagate information between graph nodes to obtain the output of GCNB. and
[0076]
[0077]
[0078] Among them, σ(·) is the activation function, W x and W′ x is the learnable weight parameter in the x-axis direction, W y and W′ y is the learnable weight parameter in the y-axis direction.
[0079] 2.1.3) Establish fusion submodule:
[0080] Reference Figure 4 , the fusion submodule includes two activation layers and one convolutional layer. After the two activation layers are connected in parallel, they are connected in series with the convolutional layer. The working process is as follows:
[0081] Two activation layers connected in parallel are used to activate the two features output by GCNB to obtain two attention weights W x and W y :
[0082]
[0083]
[0084] Multiply these two attention weights by the output of CNNB respectively. Get the initial fusion features, pass the initial fusion features through a convolution layer with a convolution kernel size of 3×3 to get the fusion features
[0085]
[0086] in, represents the initial fusion feature, ⊙ represents the dot product operation;
[0087] 2.1.4) The CNNB submodule and the GCNB submodule are connected in parallel respectively, and then connected in series with the fusion submodule to form a hybrid convolution unit, which is then connected with the other two pooling layers, that is, the first hybrid convolution unit, the first pooling layer, the second hybrid convolution unit, the second pooling layer, and the third hybrid convolution unit are cascaded in sequence to form a feature extraction module.
[0088] Through this feature extraction module, three features F with different scales containing both local information and global information are obtained 1 、F 2 、F 3 .
[0089] 2.2) Establish a progressive fusion module:
[0090] The progressive fusion module is a combination of the three features F obtained by the feature extraction module. 1 、F 2 、F 3 There are three parallel branches for fusion, each of which contains sequential sampling operations, splicing operations, and convolution operations. In the three parallel branches, F 1 、F 2 、F 3 As input, perform F 1 、F 2 、F 3 A feature fusion between 2 Taking the branch where the feature is located as an example, the process of fusion is described as follows:
[0091] 2.2.1) For features F of different sizes 1 、F 3 Perform up and down sampling to make it the same as F 2 On the same scale, the corresponding up- and down-sampling features f are obtained 1 and f 3 :
[0092] f 1 = Down(F 1 )
[0093] f 3 =Up(F3 )
[0094] Among them, Down(·) and Up(·) represent downsampling and upsampling operations respectively;
[0095] 2.2.2) Three features of the same size f 1 、F 2 and f 3 Perform channel dimension splicing to obtain the channel splicing feature f c :
[0096] f c =Concat[f 1 , F 2 , f 3 ]
[0097] =Concat[Down(F 1 ),F 2 ,Up(F 3 )]
[0098] Among them, Concat[·] represents the concatenation of channel dimensions;
[0099] 2.2.3) f c After passing through the convolution layers with kernel sizes of 1×1, 3×3, and 1×1, we get F 2 The first fusion feature F′ of the branch where the feature is located 2 :
[0100] F′ 2 =Conv 1×1,3×3,1×1 (f c )
[0101] Among them, Conv 1×1,3×3,1×1 Represents three layers of convolution with kernel sizes of 1×1, 3×3, and 1×1;
[0102] According to the process from 2.2.1) to 2.2.3), we can get the first fusion feature F′ of the other two branches. 1 and F′ 3 :
[0103] F′ 1 =Conv 1×1,3×3,1×1 (Concat[Up(F 3 ),Up(F 2 ),F 1 ])
[0104] F′ 3 =Conv 1×1,3×3,1×1 (Concat[F 3 ,Down(F 2),Down(F 1 )]);
[0105] 2.2.4) Three one-time fusion features F′ 1 , F′ 2 and F′ 3 As input, according to the process from 2.2.1) to 2.2.3), the secondary fusion features F″ of the three branches are obtained respectively. 1 , F″ 2 and F″”:
[0106] F″ 1 =Conv 1×1,3×3,1×1 (Concat[Up(F′ 3 ),Up(F′ 2 ),F′ 1 ])
[0107] F″ 2 =Conv 1×1,3×3,1×1 (Concat[Up(F′ 3 ),F′ 2 ,Down(F′ 1 )])
[0108] F″ 3 =Conv 1×1,3×3,1×1 (Concat[F′ 3 ,Down(F′ 2 ),Down(F′ 1 )]);
[0109] 2.2.5) Using three secondary fusion features F″ 1 , F″ 2 and F″ 3 As input, follow the process from 2.2.1) to 2.2.3) to get the final fusion feature F:
[0110] F=Conv 1×1,3×3,1×1 (Concat[Up(F″ 3 ),Up(F″ 2 ),F″ 1 ]).
[0111] 2.3) Establish a label update module:
[0112] Reference Figure 5 , the label update module is constructed as follows:
[0113] 2.3.1) Get the predicted label:
[0114] The output F of the progressive fusion module is convolved with two kernels of size 1×1 to obtain a two-channel feature map, and the softmax activation function is used to obtain the prediction score of the feature map. After that, the max function is applied to the prediction score in the channel dimension to obtain the prediction label.
[0115] 2.3.2) Get the prototype tag:
[0116] For the prediction scores in 2.3.1), based on the category of the predicted label, select the k=1000 feature points with the highest scores in each category, average the feature values of the k feature points in each category on the two-channel feature map, and obtain two category centers. Finally, calculate the Euclidean distance from the feature value of each feature point on the two-channel feature map to the category center, assign the category label of the category center closest to each feature point to the feature point, and obtain the prototype label of the feature point position;
[0117] 2.3.3) Update the test label, prototype label and initial label obtained in step 1 to complete the creation of the label update module:
[0118] The initial labels include three categories of labels: unchanged category labels, changed category labels, and uncertain category labels;
[0119] The predicted labels and prototype labels only include labels of unchanged categories and changed categories;
[0120] The voting label is obtained from the predicted label and the prototype label according to the voting rule. That is, at the same coordinate position, if the predicted label is the same as the prototype label, the voting label is modified to be the same as the predicted label. If the predicted label is different from the prototype label, the voting label is modified to an uncertain category label.
[0121] Use the voting labels to update the initial labels according to the update rules, that is, at the same coordinate position, replace the uncertain category labels in the initial labels with the category labels that change and do not change in the voting labels to obtain the final labels.
[0122] Step 3: Train the GECN network using three-channel input and labels.
[0123] The training of GECN is divided into two parts: pre-training and training.
[0124] The labels include initial labels used in the pre-training phase and updated initial labels used in the training phase;
[0125] 3.1) Pre-train the GECN network:
[0126] 3.1.1) Input the three-channel input constructed in 1.1) into GECN and calculate the cross entropy loss function L between the output of GECN and the initial label1 (θ):
[0127]
[0128] Among them, θ is the initial parameter of GECN obtained by random initialization, y 1 is the initial label, is the output of GECN, h θ (·) is a GECN with parameters θ, and x is the three-channel input of GECN;
[0129] 3.1.2) Use the cross entropy loss function for the GECN output and initial label, and use the gradient descent method to update the GECN parameters θ:
[0130] Assume that the initial pre-training batch m = 0, and the maximum number of iterations in the pre-training phase T = 20;
[0131] Calculate the GECN parameters θ after the current pre-training update m+1 :
[0132]
[0133] Among them, α is the learning rate in the pre-training phase, θ m are the network parameters before the current pre-training update;
[0134] 3.1.3) Repeat 3.1.2) to reach the maximum number of iterations T in the pre-training phase to complete the pre-training of GECN;
[0135] 3.2) Retrain the pre-trained GECN network:
[0136] 3.2.1) The cross entropy loss function is used for the output of the pre-trained GECN and the updated initial label. The cross entropy loss function L is calculated for the output of the pre-trained GECN and the updated initial label. 2 :
[0137] L 2 (θ')=-[y 2 logy'+(1-y 2 )log(1-y')]
[0138] Among them, θ' is the network parameter after GECN pre-training, y 2 is the updated initial label, y' is the pre-trained GECN output, y' = h θ' (x'),h θ' (·) is the GECN with parameters θ′, and x′ is the three-channel input of the GECN during the training phase;
[0139] 3.2.2) The cross entropy loss function is used for the output of the pre-trained GECN and the updated initial label, and the gradient descent method is used to update the parameters θ' of the pre-trained GECN:
[0140] Assume that the initial training batch m'=0, and the maximum number of iterations in the training phase T'=50;
[0141] Calculate the GECN parameters θ' after the current training update m+1 :
[0142]
[0143] Among them, α' is the learning rate of the training phase, θ' m are the network parameters before the current training update;
[0144] 3.2.3) Repeat 3.2.2) until the maximum number of iterations T' in the training phase is reached to obtain the trained GECN.
[0145] Step 4: Get the change result graph.
[0146] 4.1) Input the tested SAR image into the trained GECN to obtain the GECN output. Use the max function on the GECN output in the channel dimension to obtain the channel with the maximum value:
[0147]
[0148] Among them, x 0 and x 1 They are the values of channel 0 and channel 1 of the GECN output features respectively;
[0149] 4.2) According to the channel where the maximum value of each position on the GECN output feature is located, which corresponds to the attribute of the category of the pixel sample at that position, the category of each pixel sample can be obtained, that is, Channel = 0 represents that the pixel sample belongs to the unchanged category, and Channel = 1 represents that the pixel sample belongs to the changed category, thus completing the change detection of the SAR image.
[0150] The effect of the present invention can be further illustrated by the following simulation:
[0151] 1. Simulation conditions
[0152] The simulation environment of the present invention selected the framework of python 3.8+pytorch 1.7 and was completed on a workstation with GeForce RTX2080Ti and 11G memory.
[0153] The four datasets used in the simulation are the Bern dataset, the Ottawa dataset, the Yellow River I dataset, and the Yellow River IV dataset.
[0154] The Bern dataset was captured by the European Remote Sensing 2 satellite sensor in April and May 1999, with an image size of 301×301;
[0155] The Ottawa dataset was captured by radar satellite sensors in May and August 1997, with image sizes of 290 × 350;
[0156] The Yellow River I dataset and the Yellow River IV dataset are two classic areas of the Yellow River dataset, captured by the Radarsat-2 satellite sensor in June 2008 and June 2009, respectively, with image sizes of 289×257 and 306×291.
[0157] 2. Simulation content
[0158] Under the above simulation conditions, the present invention and four existing methods, DDNet, SAFNet, MsCapsNet, and MSDC, are used to perform change detection simulations on four data sets. The visualization results are shown in Figure 2. Figure 6 The numerical results are shown in Tables 1 to 4. Among them:
[0159] Figure 6 (a) is the pre-time image of the four datasets;
[0160] Figure 6 (b) is the post-time image of the four datasets;
[0161] Figure 6 (c) is the ground truth of the four datasets;
[0162] Figure 6 (d) is the visualization result of the existing method DDNet on four datasets;
[0163] Figure 6 (e) is the visualization result of the existing method SAFNet on four datasets;
[0164] Figure 6 (f) is the visualization result of the existing method MsCapsNet on four datasets;
[0165] Figure 6 (g) is the visualization result of the existing method MSDC on four datasets;
[0166] Figure 6(h) is the visualization result of the present invention on four data sets;
[0167] from Figure 6 It can be seen that the present invention can effectively reduce the generation of noise points and detect the boundaries of changing objects more accurately, which fully demonstrates the superiority of the method proposed in the present invention.
[0168] In order to quantitatively illustrate the performance of the proposed method, numerical performance indicators commonly used in change detection tasks are selected to measure the differences between the above-mentioned existing methods and the present invention. The performance indicators include overall error OE, correct classification percentage PCC and Kappa coefficient KC. The results are shown in the following table, where:
[0169] Table 1 shows the numerical results of the present invention and four existing methods on the Bern dataset;
[0170] Table 2 shows the numerical results of the present invention and four existing methods on the Ottawa dataset;
[0171] Table 3 is the numerical results of the present invention and four existing methods on the Yellow River I dataset;
[0172] Table 4 is the numerical results of the present invention and four existing methods on the Yellow Rive IV data set;
[0173] Table 1 Results of Bern dataset
[0174]
[0175] Table 2 Results of Ottawa dataset
[0176]
[0177] Table 3 Results of Yellow Rive I dataset
[0178]
[0179] Table 4 Results of Yellow Rive IV dataset
[0180]
[0181] It can be seen from Table 1 to Table 4 that the overall error of the present invention is smaller and the change detection accuracy is higher, which further illustrates the superiority of the method proposed in the present invention.
[0182] The sources of the four prior arts are:
[0183] DDNet is a method for SAR image change detection published by Qu Xiaofan et al. in IEEE, namely: X.Qu, F.Gao, J.Dong, Q.Du, and H.-C.Li, "Change detection in synthetic aperture radar images using adual-domain network," IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2021;
[0184] SAFNet is a method for SAR image change detection published by Yunhao Gao et al. in IEEE, namely: Y.Gao, F.Gao, J.Dong, Q.Du, and H.-C.Li, "Synthetic aperture radar image changedetection via siamese adaptive fusion network," IEEE Journal of SelectedTopics in Applied Earth Observations and Remote Sensing, vol.14, pp.10 748–10760, 2021;
[0185] MsCapsNet is a method for SAR image change detection published by Gao Yunhao et al. in IEEE, namely: Y.Gao, F.Gao, J.Dong, and H.-C.Li, "SAR image change detection based on multiscalecapsule network," IEEE Geoscience and Remote Sensing Letters, vol.18, no.3, pp.484–488, 2020;
[0186] MSDC is a method for change detection in SAR images published by Dong Huihui et al. in IEEE, namely: Dong H, Ma W, Jiao L, et al. A multiscale self-attention deep clustering for change detection in SAR images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 1-16.
Claims
1. A SAR image change detection method based on hybrid convolution, characterized in that: The steps include: (1) Construct three-channel input and initial labels: 1a) A logarithmic ratio operator is applied to the dual-time SAR image to obtain a difference image, and the difference image is concatenated with the dual-time image in the channel dimension to obtain a three-channel input; 1b) Cluster the difference graph in 1a) using a fuzzy clustering algorithm to obtain an initial label; (2) Build a graph convolution enhanced convolutional network GECN: 2a) establishing a feature extraction module consisting of a convolutional neural network submodule CNNB, a graph convolutional neural network submodule GCNB and a fusion submodule, wherein CNNB and GCNB are connected in parallel and then in series with the fusion submodule; 2b) establishing a progressive fusion module that sequentially performs up-sampling and down-sampling operations, concatenation operations, and convolution operations; 2c) Establishing a label update module that updates the training labels by voting with the predicted labels and prototype labels; 2d) connecting the feature extraction module, the progressive fusion module and the label update module in series to form a graph convolution enhanced convolutional network GECN; (3) Use three-channel input and labels to train GECN: 3a) In the initial 20 training batches, the three-channel input is fed into GECN, the cross entropy loss function is applied to the GECN output and the initial label, and the GECN parameters are updated using the gradient descent method to complete the GECN pre-training; 3b) In the subsequent 50 training batches, the three-channel input is fed into the pre-trained GECN, the initial label is updated using the label update module, the cross entropy loss function is applied to the output of the pre-trained GECN and the updated initial label, and the gradient descent method is used to further update the pre-trained GECN parameters to complete the training of the GECN; (4) The test SAR image is input into the trained GECN to obtain the GECN output. The max function is used on the GECN output to obtain the category of each pixel sample on the SAR image, which is the final change detection result of the SAR image.
2. The method according to claim 1, characterized in that In step 1b), the difference graph is clustered using a fuzzy clustering algorithm, which is implemented as follows: 1b1) Determine the number c of fuzzy clustering categories and the number of iterations k; 1b2) Initialize the membership matrix u with random numbers between 0 and 1 so that its sum is 1. The formula is: Among them, u ij is the membership of the jth pixel sample to the i-th category, and n is the number of pixel samples; 1b3) Calculate the cluster center based on the membership degree: Among them, c i is the i-th cluster center, x j is the j-th pixel sample; 1b4) Calculate the sum of squared errors J from pixel samples to various center points: 1b5) Find the extreme value of formula 1b4) and when the sum of squared errors is minimized, we get the updated membership matrix u ij ; 1b6) Repeat 1b3) to 1b5) until the iteration number k is reached and the initial label is obtained.
3. The method according to claim 1, characterized in that The graph convolution neural network submodule GCNB in step 2a) includes two parallel pooling layers and four graph convolution layers connected in sequence: The two parallel pooling layers are used to pool the input of GCNB along the x-axis and y-axis directions respectively to compress the input features to obtain graph nodes; The four graph convolutional layers are connected in series and then in parallel to transfer the relationship between graph nodes to obtain two output features.
4. The method according to claim 1, characterized in that: The fusion submodule in step 2a) includes two parallel-connected activation layers and a convolutional layer connected in sequence: The two activation layers connected in parallel are used to first activate the two features output by GCNB to obtain attention weights, and then the two attention weights are respectively multiplied by the output of CNNB to obtain fusion features; This convolution layer is used to further mine the information of fused features and obtain the output features of the fused sub-module.
5. The method according to claim 1, characterized in that In step 3a), the cross entropy loss function is applied to the GECN output and the initial label, and the GECN parameters are updated using the gradient descent method, which is implemented as follows: 3a1) Calculate the cross entropy loss function L1 between the output of GECN and the initial label: Among them, y1 is the initial label, is the output of GECN, h θ (·) is a GECN with parameters θ, and x is the three-channel input of GECN; 3a2) Update GECN parameters θ using gradient descent: Assume that the initial pre-training batch m = 0, θ0 is the initial parameter obtained by random initialization, and the maximum number of iterations in the pre-training stage is T = 20; Calculate the GECN parameters θ after the current pre-training update m+1 : Among them, α is the learning rate; 3a3) Repeat 3a2) to reach the maximum number of iterations T in the pre-training phase to complete the pre-training of GECN.
6. The method according to claim 1, characterized in that In step 3b), the cross entropy loss function is applied to the output of the pre-trained GECN and the updated initial label, and the pre-trained GECN parameters are updated using the gradient descent method, which is implemented as follows: 3b1) Calculate the cross entropy loss function L2 between the output of the pre-trained GECN and the updated initial label: L2(θ')=-[y2log y'+(1-y2)log(1-y')] Among them, y2 is the updated initial label, y' is the GECN output after pre-training, y' = h θ' (x'),h θ' (·) is the GECN with parameters θ′, and x′ is the three-channel input of the GECN during the training phase; 3b2) Update the parameters θ' of the pre-trained GECN using gradient descent: Assume that the initial training batch m'=0, θ'0 is the parameter obtained after pre-training GECN, and the maximum number of iterations in the training phase is T'=50; Calculate the GECN parameters θ' after the current training update m+1 : 3b3) Repeat 3b2) until the maximum number of iterations T' in the training phase is reached, and the trained GECN is obtained.
Citation Information
Patent Citations
SAR image classification algorithm combining graph convolutional network and Markov random field
CN113486967A
Progressive bone age assessment method based on multi-granularity feature fusion
CN114049303A