Semantic change detection method for high-resolution remote sensing images based on self-attention feature fusion

By adopting the method of self-attention feature fusion in semantic change detection, combining a combined deep neural network and a multi-task decoding layer, the problems of poor detection accuracy of land object change types and model overfitting in the existing technology are solved, and higher detection accuracy and generalization capabilities are achieved.

CN116486255BActive Publication Date: 2025-05-16FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310253842.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-16
Publication Date
2025-05-16
Estimated Expiration
2043-03-16

AI Technical Summary

Technical Problem

When detecting the type of change of land objects, the existing semantic change detection methods tend to ignore the correlation between the categories before and after the changes of land objects, and due to the imbalance of the transition categories, poor detection accuracy and model overfitting.

Method used

The semantic change detection method of self-attention feature fusion of high-resolution remote sensing image is adopted, and the recognition accuracy of changing regions and landform types is improved through a combined deep neural network, attention mechanism and multi-task decoding layer, and the overfitting problem of single-task model is overcome.

Benefits of technology

It effectively improves the accuracy and generalization ability of land use/land coverage type detection, reduces missed inspections and missed inspections, and provides important technical support to government management departments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486255B_ABST
    Figure CN116486255B_ABST
Patent Text Reader

Abstract

The present invention relates to a semantic change detection method for high-resolution remote sensing images with self-attention feature fusion. The method combines deep learning technology to construct a semantic change detection model for land use / land cover (LULC) of high-resolution remote sensing images. The model uses a convolutional neural network CNN and a Transformer combined encoding layer to extract deep features of dual-phase high-resolution images for preprocessed and cropped high-resolution remote sensing images, and cooperates with a self-attention feature fusion module and a multi-task deconvolution decoding layer to perform LULC change area detection and classification of land object types in the previous and next phases within the change area. The present invention combines the idea of ​​multi-task learning, integrates change area detection and change type recognition into one, and realizes automated dual-phase high-resolution image LULC semantic change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research field of land use / land cover change detection of dual-temporal high-resolution remote sensing images, and in particular to a high-resolution remote sensing image semantic change detection method based on self-attention feature fusion. Background Art

[0002] Human activities affect the surface environment to varying degrees. Remote sensing image change detection technology can effectively record various changes in surface land use / land cover (LULC), which is of great significance to urban and rural planning, disaster monitoring, natural resource management and ecological protection. With the development of science and technology, more refined LULC change detection results can be obtained by combining high-resolution images and deep learning technology. However, most of the existing change detection methods focus on binary change area detection of multi-temporal remote sensing images (i.e., only focusing on changes and no changes between images). In current practical applications, accurately detecting the type of ground object change (i.e., what land use type the changed ground object changes from) has become one of the important needs.

[0003] Semantic Change Detection (SCD) is a pixel-level change detection method that can obtain more information about surface changes. It can detect the change area and identify the type of objects in the area before and after the change. Usually, semantic change detection using deep learning technology only needs to input dual-temporal images and then use a neural network model to obtain the change type detection result. However, using a single task to directly output the change type often ignores the category correlation before and after the change of the object, and the output result will lead to poor final accuracy and model overfitting due to the imbalance of the change category.

[0004] Multi-task learning is a machine learning method that puts multiple related tasks together for learning. Several related subtasks are learned in parallel while sharing weight parameters, and then the correlation between tasks and the overall generalization of the model are improved through gradient back propagation. Multi-task learning has been applied in medical image segmentation and real scene target detection, and in the field of remote sensing, multi-task learning has also been applied to research such as plot extraction, semantic segmentation and urban classification and has achieved good results, but its application in change detection is relatively small. Based on this, the present invention applies multi-task learning to semantic change detection of LULC and constructs a high-resolution remote sensing image semantic change detection method with self-attention feature fusion. Summary of the invention

[0005] The purpose of the present invention is to provide a high-resolution remote sensing image semantic change detection method with self-attention feature fusion, which flexibly uses a combined deep neural network, an attention mechanism and a multi-task decoding layer to overcome the problems of low accuracy in identifying the category of object changes in single-task semantic change detection and easy overfitting of the model, and effectively improves the accuracy and generalization ability of LULC change type detection.

[0006] To achieve the above object, the technical solution of the present invention is: a high-resolution remote sensing image semantic change detection method based on self-attention feature fusion, comprising the following steps:

[0007] Step S1, obtaining two high-resolution remote sensing images of the same area at different phases, and preprocessing the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration and image resampling operations;

[0008] Step S2, cropping and data augmentation processing are performed on the pre-processed images, and a high-resolution image land use / land cover LULC tile sample dataset is constructed through manual screening and correction;

[0009] Step S3, based on the LULC tile sample data set constructed in step S2, the corresponding dual-temporal image sample pairs are input into two feature coding layers respectively to identify and extract the deep feature information of the dual-temporal image;

[0010] Step S4, constructing a self-attention feature fusion module, using the self-attention feature fusion module to fuse the deep feature information of the dual-phase image extracted in step S3 to obtain fused features;

[0011] Step S5, constructing a multi-task feature decoding layer for the fusion features obtained in step S4; in the multi-task decoding layer, one task is used for detecting the changed area, and the other task is used for identifying the type of ground objects in the LULC feature information of the previous and next phases in the changed area, and the changed area prediction result is masked to the change type result to realize spatial area constraint;

[0012] Step S6: Combined with the corresponding real tile sample labels, use the multi-task loss function to calculate the difference metric of the model prediction results and select the parameter optimizer to update and optimize the parameters of the model back propagation. Through multiple iterative training, adjust the model parameters to continuously reduce the distance between the model prediction results and the real labels so that the prediction results are more consistent with the real labels, thereby improving the model detection capability.

[0013] Step S7: Using the model optimized in step S6, complete the identification of LULC change types in the entire region.

[0014] In one embodiment of the present invention, step S3 is specifically implemented as follows:

[0015] Step S31, fusion CNN and Transformer combination as the feature encoding layer; wherein CNN uses the residual neural network ResNet34 to extract shallow features, and the specific steps include: (1) using a single convolution layer with a convolution kernel size of 7×7 and a maximum pooling layer with a convolution kernel size of 3×3 to extract the first layer features of the initial image; (2) using the residual block as the basic module to extract the second layer features, which is divided into two parts: the first part selects 3 basic modules and 1 pooling layer to extract features; the second part selects 4 basic modules and 1 pooling layer to extract features; (3) using dimensional matrix conversion to reorganize the dimensions of the image feature channel C, length H and width W, that is, the image dimension is changed from (1, C, H, W) to (1, H×W, C);

[0016] Step S32, using Swin-Transformer with sliding window image processing and hierarchical design to perform deep feature extraction on the shallow features obtained in step S32, the specific steps include: (1) using 10 Swin-Transformer Block basic modules and 3 block dimensionality reduction sampling modules to extract the third layer features of the reorganized image, which is divided into three parts: the first part selects 2 basic modules and 1 block dimensionality reduction sampling module; the second part selects 6 basic modules and 1 block dimensionality reduction sampling module; the third part selects 2 basic modules and 1 block dimensionality reduction sampling module; (2) using dimensional matrix transformation to perform anti-reorganization of the image feature channel C, length H and width W, that is, the image dimension is transformed from (1, H×W, C) to (1, C, H, W).

[0017] In one embodiment of the present invention, step S4 is specifically implemented as follows:

[0018] Step S41, using two feature coding layers with shared weights to obtain deep features of dual-phase images;

[0019] Step S42, performing feature channel dimension combination and convolution on the deep features of the dual-phase image obtained in step S41, and then using a self-attention module to enhance the internal correlation of the feature mapping;

[0020] Step S43: While performing step S42, the deep features of the dual-phase images obtained in step S41 are subjected to the feature difference method to obtain the differential fusion results of the deep features of the dual-phase images. Then, the self-attention module is used to enhance the internal correlation of the feature mapping, and the result is spatially superimposed with the result of step S42.

[0021] In one embodiment of the present invention, step S5 is specifically implemented as follows:

[0022] Step S51, constructing a multi-task branch upsampling layer, which includes two tasks: a change region detection branch and a change type recognition branch, and the two tasks adopt a weight sharing method to achieve parallel optimization of the change region detection and change type recognition accuracy;

[0023] Step S52: for the branch task of changed area detection, three deconvolution layers with a convolution kernel size of 3×3 and three upsampling layers of bilinear interpolation method are selected to form a decoder for changed area detection;

[0024] Step S53: For the change type recognition branch task, two deconvolution layers with a convolution kernel size of 3×3, two deconvolution layers with a convolution kernel size of 1×1, and three upsampling layers of bilinear interpolation are selected to form a decoder for LULC change type detection; the dual-branch parallel decoder is unified into a single-branch decoder, and only two deconvolution layers are used in the last layer of the decoder to output the ground feature type recognition results before and after the time phase in the change area, so as to complete the detection of semantic change information;

[0025] Step S54: Mask the multi-classification detection result of the semantic change information in step S53 with the binary classification detection result of the changed area in step S52 to achieve spatial constraints on the prediction results and further improve the overall accuracy of semantic change detection.

[0026] In one embodiment of the present invention, step S6 is specifically implemented as follows:

[0027] Step S61, based on the multi-task decoder used by the multi-task feature decoding layer in step S5, two loss functions are used to calculate and the final calculation results are added as the final loss function value of the model; the loss functions used include Focal Loss and binary cross entropy loss function BCE Loss, wherein Focal Loss is used for loss calculation of change type detection, and BCE Loss is used for loss calculation of change area detection;

[0028] Focal Loss is defined as follows:

[0029] L Focal =(-α t (1-q(x)) γ )·(∑(p(x)·log(q(x))+(1-p(x))·log(1-q(x))

[0030] In the formula, p(x) represents the predicted probability distribution, q(x) represents the true probability distribution, and α t Represents the weight of positive and negative samples, γ is the gamma coefficient, which is used to reduce the loss calculation result value of simple class samples;

[0031] BCE Loss is defined as follows:

[0032] L BCE =-p(x)·log((q(x))+(1-p(x))·log(1-q(x))

[0033] In the formula, p(x) represents the predicted probability distribution, and q(x) represents the true probability distribution;

[0034] Step S62, the parameter optimizer selects the iterative momentum stochastic gradient descent optimizer SGDM, and the relevant parameters are set as: learning rate 0.01, weight decay value 0.0001, and iterative momentum value 0.9; the parameter optimizer updates and optimizes the model parameters through back propagation after completing the loss function to calculate the difference metric, and achieves model fitting while reducing the loss function through multiple iterative training.

[0035] Compared with the prior art, the present invention has the following beneficial effects: (1) It uses a multi-task constraint approach to improve the generalization of model detection, overcoming the problem of easy overfitting of traditional single-task semantic change detection models; (2) It uses a combined deep neural network and attention mechanism to improve the accuracy of detection of changed areas and identification of changed ground object type information, overcoming the problems of missed detection and false detection that are prone to occur in traditional methods, and further improving the accuracy of LULC ground object change type detection in high-resolution remote sensing images, providing important methodological and technical support for government management departments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 The figure is a schematic diagram of a method flow of an embodiment of the present invention.

[0037] Figure 2 This is a structural diagram of a combined feature encoder according to an embodiment of the present invention.

[0038] Figure 3 This is a structural diagram of the self-attention feature fusion module of an embodiment of the present invention.

[0039] Figure 4 This is a graph showing the semantic change detection results of part of the SECOND dataset according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0041] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0043] like Figure 1 As shown, this embodiment provides a high-resolution remote sensing image semantic change detection method based on self-attention feature fusion, including the following steps:

[0044] Step S1, obtaining two high-resolution remote sensing images of the same area at different phases, and preprocessing the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration, and image resampling;

[0045] Step S2, cropping and data augmentation processing are performed on the preprocessed images, and a high-resolution image land use / land cover (LULC) tile sample dataset is constructed through appropriate manual screening and correction;

[0046] Step S3, based on the LULC tile sample data set constructed in step S2, the corresponding dual-temporal image sample pairs are input into two feature coding layers respectively to identify and extract the deep feature information of the dual-temporal image;

[0047] Step S4: construct a self-attention feature fusion module based on the advantages and disadvantages of the channel dimension joint method and the feature difference method used in the traditional feature fusion, and use the self-attention feature fusion module to fuse the features of the dual-phase image LULC feature information extracted in step S3;

[0048] Step S5, construct a multi-task feature decoding layer for the fusion features obtained in step S4. In the multi-task decoding layer, one task is used to detect the change area, and the other task is used to identify the LULC feature types before and after the change area. The change area prediction result is masked to the change type result to realize spatial area constraint;

[0049] Step S6: Combined with the corresponding real tile sample labels, use the multi-task loss function to calculate the difference metric of the model prediction results and select the parameter optimizer to update and optimize the parameters of the model back propagation. Through multiple iterative training, adjust the model parameters to continuously reduce the distance between the model prediction results and the real labels, so that the prediction results are more consistent with the real labels, thereby improving the model detection capability.

[0050] Step S7: Using the above parameter-optimized model, complete the identification of LULC change types in the entire region.

[0051] In this embodiment, step S3 specifically includes the following steps:

[0052] Step S31: Based on the advantages and disadvantages of the convolutional neural network (CNN) and the Transformer model, CNN is integrated.

[0053] and Transformer as the feature encoding layer. Among them, CNN uses the residual neural network ResNet34 for shallow feature extraction. The specific steps include: (1) using a single convolution layer with a convolution kernel size of 7×7 and a maximum pooling layer with a convolution kernel size of 3×3 to extract the first layer features of the initial image; (2) using the residual block as the basic module to extract the second layer features, this layer is divided into two parts: the first part selects 3 basic modules and 1 pooling layer to extract features; the second part selects 4 basic modules and 1 pooling layer to extract features; (3) using dimensional matrix transformation to reorganize the dimensions of the image feature channel C, length H and width W. Specifically, the image dimension is transformed from (1, C, H, W) to (1, H×W, C).

[0054] Step S32, using Swin-Transformer with sliding window image processing and hierarchical design to perform deep feature extraction on the shallow features obtained in step S32, the specific steps include: (1) using 10 Swin-Transformer Block basic modules and 3 block dimensionality reduction sampling modules to extract the third layer features of the reorganized image, the layer is divided into three parts: the first part selects 2 basic modules and 1 block dimensionality reduction sampling module; the second part selects 6 basic modules and 1 block dimensionality reduction sampling module; the third part selects 2 basic modules and 1 block dimensionality reduction sampling module; (2) using dimensional matrix transformation to perform anti-reorganization of the image feature channel C, length H and width W. Specifically, the image dimension is transformed from (1, H×W, C) to (1, C, H, W).

[0055] In this embodiment, step S4 specifically includes the following steps:

[0056] Step S41, using two feature encoders with shared weights to obtain deep features of dual-phase images;

[0057] Step S42, performing feature channel dimension combination and convolution on the dual-phase deep features obtained in step S41, and then using a self-attention module to enhance the internal correlation of feature mapping;

[0058] Step S43: While performing step S42, the dual-phase deep features obtained in step S41 are subjected to the feature difference method to obtain the differential fusion result of the dual-phase features, and then the self-attention module is used to enhance the internal correlation of the feature mapping, and the result is spatially superimposed with the result of step S42.

[0059] In this embodiment, step S5 specifically includes the following steps:

[0060] Step S51, constructing a multi-task branch upsampling layer, which includes two tasks: a change region detection branch and a change type recognition branch, and the two tasks adopt a weight sharing method to achieve parallel optimization of the change region detection and change type recognition accuracy;

[0061] Step S52: for the branch task of changed area detection, three deconvolution layers with a convolution kernel size of 3×3 and three upsampling layers of bilinear interpolation method are selected to form a decoder for changed area detection;

[0062] Step S53: For the change type recognition branch task, two deconvolution layers with a convolution kernel size of 3×3, two deconvolution layers with a convolution kernel size of 1×1, and three upsampling layers of bilinear interpolation are selected to form a decoder for LULC change type detection; in order to reduce the amount of calculation of the model when interpreting deep image features, the original dual-branch parallel decoder is unified into a single-branch decoder, and only two deconvolution layers are used in the last layer of the decoder to output the ground feature type recognition results before and after the time phase in the change area respectively, so as to complete the detection of semantic change information;

[0063] Step S54: Mask the multi-classification detection result of the semantic change information in step S53 with the binary classification detection result of the changed area in step S52 to achieve spatial constraints on the prediction results and further improve the overall accuracy of semantic change detection.

[0064] In this embodiment, step S6 specifically includes the following steps:

[0065] Step S61, the purpose of the loss function is to calculate and evaluate the degree of difference between the model prediction result and the actual situation. Based on the multi-task decoder used in step S5, two loss functions are used to calculate and add the final calculation results as the final loss function value of the model. The loss functions used include Focal Loss and Binary Cross Entropy Loss (BCE Loss), where Focal Loss is used for loss calculation of change type detection and BCE Loss is used for loss calculation of change area detection.

[0066] Focal Loss is defined as follows:

[0067] L Focal =(-α t (1-q(x)) γ )·(∑(p(x)·log(q(x))+(1-p(x)·log(1-q(x))#

[0068] In the formula, p(x) represents the predicted probability distribution, q(x) represents the true probability distribution, and α t Represents the weight of positive and negative samples, and γ is the gamma coefficient (used to reduce the loss calculation result value of simple class samples).

[0069] BCE Loss is defined as follows:

[0070] L BCE =-p(x)·log((q(x))+(1-p(x))·log(1-q(x))

[0071] Where p(x) represents the predicted probability distribution and q(x) represents the true probability distribution.

[0072] Step S62, the purpose of the parameter optimizer is to update and optimize the model parameters through back propagation after the loss function calculates the difference metric, and achieves model fitting while reducing the loss function through multiple iterative training. This method selects the iterative momentum stochastic gradient descent optimizer (Stochastic Gradient Descent with momentum, SGDM), and the relevant parameters are set as follows: learning rate (Learning Rate) 0.01, weight decay value (WeightDecay) 0.0001, iterative momentum value (Momentum) 0.9.

[0073] Compared with the existing methods, the present invention has the following advantages: (1) It uses multi-task constraints to improve the generalization of model detection, overcoming the problem of easy overfitting of traditional single-task semantic change detection models; (2) It uses a combined deep neural network and attention mechanism to improve the accuracy of detection of changed areas and identification of changed ground object type information, overcoming the problems of missed detection and false detection that are prone to occur in traditional methods, further improving the accuracy of LULC ground object change type detection in high-resolution remote sensing images, and providing important methodological and technical support for government management departments.

[0074] In this embodiment, the SECOND semantic change detection dataset after preprocessing and cropping is used, with a spatial resolution of 0.5-2m and three bands of red, green and blue. This embodiment uses 2962 tile images with a size of 512×512 after tile cropping and corresponding labels for model training and verification.

[0075] like Figure 2This is a structural diagram of the combined deep learning model of this embodiment. The feature encoder of the model consists of 6 feature extraction layers. The structure of the first three layers adopts the first three layers of ResNet34 Backbone. The first layer uses a convolution layer with a convolution kernel size of 7×7 and a maximum pooling layer to extract features. The second and third layers both use a residual block (Residualblock) composed of two convolution layers with a convolution kernel size of 3×3, a residual convolution layer with a convolution kernel size of 1×1, and a maximum pooling layer; and the structure of the last three layers all adopts the basic component modules of Swin-Transformer, which are divided into two parts: the first part consists of a fully connected layer, a fixed window multi-head attention, and a multi-layer perceptron, and the second part consists of a fully connected layer, a sliding window multi-head attention, and multiple perceptrons. The feature decoder of the model has 4 layers, each of which is composed of a deconvolution layer with a convolution kernel size of 1×1 and a bilinear interpolation upsampling layer. The same decoder is used to identify the type of features before and after the change area, and only the last layer is branched to realize the predicted output of the type of changed features before and after.

[0076] like Figure 3 This is a structural diagram of the self-attention feature fusion module used in this embodiment. The module first uses the channel combination method and the feature difference method to perform feature fusion on the dual-phase image, and uses the self-attention mechanism to enhance the correlation of the internal data feature information of each result. Finally, spatial superposition is used to achieve the fusion of the results obtained by the two methods.

[0077] like Figure 4 The following is a partial experimental result diagram of the SECOND semantic change detection dataset after processing in this embodiment. It can be seen from the figure that the semantic change detection prediction results obtained by the method of the present invention have a high degree of consistency compared with the true label; the detection accuracy of the change area between the previous and next phase images and the refinement of the change shape are both high, with less missed detection and false detection; the identification of the type of objects in the change area of ​​the previous and next phase images is relatively accurate; the experimental results show the effectiveness of the method in land use / land cover semantic change detection.

[0078] Compared with the existing methods, the high-resolution remote sensing image semantic change detection model of multi-task learning in the present invention uses multi-task constraints to improve the generalization of model detection, overcoming the problem of easy overfitting of traditional single-task semantic change detection models; at the same time, a combined deep neural network and attention mechanism are used to improve the accuracy of detection of changed areas and identification of changed land object type information, overcoming the problems of missed detection and false detection that are prone to occur in traditional methods, and further improving the accuracy of remote sensing land use / land cover land object change type detection.

[0079] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.

Claims

1. A high-resolution remote sensing image semantic change detection method based on self-attention feature fusion, characterized in that: The steps include: Step S1, obtaining two high-resolution remote sensing images of the same area at different phases, and preprocessing the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration and image resampling operations; Step S2, cropping and data augmentation processing are performed on the pre-processed images, and a high-resolution image land use and land cover LULC tile sample dataset is constructed through manual screening and correction; Step S3, based on the LULC tile sample data set constructed in step S2, the corresponding dual-temporal image sample pairs are input into two feature coding layers respectively to identify and extract the deep feature information of the dual-temporal image; Step S4, constructing a self-attention feature fusion module, using the self-attention feature fusion module to fuse the deep feature information of the dual-phase image extracted in step S3 to obtain fused features; Step S5, construct a multi-task feature decoding layer for the fusion features obtained in step S4; in the multi-task decoding layer, one task is used for detecting the change area, and the other task is used for identifying the type of ground objects in the LULC feature information of the previous and next phases within the change area, and the change area prediction result is masked to the change type result to realize spatial area constraint; the specific implementation is as follows: Step S51, constructing a multi-task branch upsampling layer, which includes two tasks: a change region detection branch and a change type recognition branch, and the two tasks adopt a weight sharing method to achieve parallel optimization of the change region detection and change type recognition accuracy; Step S52: for the branch task of changed area detection, three deconvolution layers with a convolution kernel size of 3×3 and three upsampling layers of bilinear interpolation method are selected to form a decoder for changed area detection; Step S53: for the change type recognition branch task, two deconvolution layers with a convolution kernel size of 3×3, two deconvolution layers with a convolution kernel size of 1×1, and three upsampling layers of bilinear interpolation method are selected to form a decoder for LULC change type detection; The dual-branch parallel decoder is unified into a single-branch decoder, and only two deconvolution layers are used in the last layer of the decoder to output the object type recognition results before and after the change area, so as to complete the detection of semantic change information. Step S54, masking the multi-classification detection result of the semantic change information in step S53 with the binary classification detection result of the change area in step S52, so as to realize the spatial constraint of the prediction result; Step S6: Combined with the corresponding real tile sample labels, use the multi-task loss function to calculate the difference metric of the model prediction results and select the parameter optimizer to update and optimize the parameters of the model back propagation, and adjust the model parameters through multiple iterative training; Step S6 comprises: Step S61, based on the multi-task decoder used by the multi-task feature decoding layer in step S5, two loss functions are used to calculate and the final calculation results are added as the final loss function value of the model; the loss functions used include FocalLoss and the binary cross entropy loss function BCE Loss, wherein Focal Loss is used for loss calculation of change type detection, and BCE Loss is used for loss calculation of change area detection; Focal Loss is defined as follows: L Focal =(-α t (1-q(x)) γ )·(∑(p(x)·log(q(x)+(1-p(x))·log(1-q(x)) In the formula, p(x) represents the predicted probability distribution, q(x) represents the true probability distribution, and α t Represents the weight of positive and negative samples, γ is the gamma coefficient, which is used to reduce the loss calculation result value of simple class samples; BCE Loss is defined as follows: L BCE =-p(x)·log((q(x)+(1-p(x))·log(1-q(x)) In the formula, p(x) represents the predicted probability distribution, and q(x) represents the true probability distribution; Step S7: Using the model optimized in step S6, complete the identification of LULC change types in the entire region.

2. The high-resolution remote sensing image semantic change detection method based on self-attention feature fusion according to claim 1 is characterized in that: The step S3 is specifically implemented as follows: Step S31, fusion CNN and Transformer combination as the feature encoding layer; wherein CNN uses the residual neural network ResNet34 to extract shallow features, and the specific steps include: (1) using a single convolution layer with a convolution kernel size of 7×7 and a maximum pooling layer with a convolution kernel size of 3×3 to extract the first layer features of the initial image; (2) using the residual block as the basic module to extract the second layer features, which is divided into two parts: the first part selects 3 basic modules and 1 pooling layer to extract features; the second part selects 4 basic modules and 1 pooling layer to extract features; (3) using dimensional matrix conversion to reorganize the dimensions of the image feature channel C, length H and width W, that is, the image dimension is changed from (1, C, H, W) to (1, H×W, C); Step S32, using Swin-Transformer with sliding window image processing and hierarchical design to perform deep feature extraction on the shallow features obtained in step S32, the specific steps include: (1) using 10 Swin-Transformer Block basic modules and 3 block dimensionality reduction sampling modules to extract the third layer features of the reorganized image, which is divided into three parts: the first part selects 2 basic modules and 1 block dimensionality reduction sampling module; the second part selects 6 basic modules and 1 block dimensionality reduction sampling module; the third part selects 2 basic modules and 1 block dimensionality reduction sampling module; (2) using dimensional matrix transformation to perform anti-reorganization of the image feature channel C, length H and width W, that is, the image dimension is transformed from (1, H×W, C) to (1, C, H, W).

3. The high-resolution remote sensing image semantic change detection method based on self-attention feature fusion according to claim 1, characterized in that: The step S4 is specifically implemented as follows: Step S41, using two feature coding layers with shared weights to obtain deep features of dual-phase images; Step S42, performing feature channel dimension combination and convolution on the deep features of the dual-phase image obtained in step S41, and then using a self-attention module to enhance the internal correlation of the feature mapping; Step S43: While performing step S42, the deep features of the dual-phase images obtained in step S41 are subjected to the feature difference method to obtain the differential fusion results of the deep features of the dual-phase images. Then, the self-attention module is used to enhance the internal correlation of the feature mapping, and the differential fusion results are spatially superimposed with the results of step S42.

4. The high-resolution remote sensing image semantic change detection method based on self-attention feature fusion according to claim 1, characterized in that: The step S6 further includes: Step S62, the parameter optimizer selects the iterative momentum stochastic gradient descent optimizer SGDM, and the relevant parameters are set as: learning rate 0.01, weight decay value 0.0001, and iterative momentum value 0.9; the parameter optimizer updates and optimizes the model parameters through back propagation after completing the loss function to calculate the difference metric, and achieves model fitting while reducing the loss function through multiple iterative training.

Citation Information

Patent Citations

  • Remote sensing image semantic change detection method based on twin Transformers

    CN114842351A

  • Transform and dense feature fusion-based remote sensing image change detection method and system

    CN115690002A