An efficient remote sensing image change detection method based on semantic guided hybrid network
Patent Information
- Application Number
- CN202610011496.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-01-06
AI Technical Summary
中国专利公开号CN119338780A,发明名称是一种基于Transformer和图语义引导的遥感图像变化检测方法及系统,该专利公开了图像处理技术领域的一种基于Transformer和图语义引导的遥感图像变化检测方法及系统,其不足之处是现有技术在遥感图像变化检测中难以有效捕获像素间远距离的依赖关系,且直接相减的特征融合方式容易破坏特征结构并产生噪音干扰;由上述内容可知,当前采用CNN或Transformer的算法仍然存在计算复杂度高、跨尺度和双时相像信息交互不足等问题,常见的多尺度融合方法是通过拼接或相加的方式直接融合浅层和深层特征,该方法虽然简便,但适应性较差,缺乏针对双时相图像的时空交互机制,限制了对复杂变化场景的判别能力
[0092]1)本发明提出的一种基于语义引导混合网络的高效遥感图像变化检测方法,针对现有混合架构高昂的计算开销,在编码阶段降低和
的空间分辨率,最小化计算开销,利用交叉注意力,强制双时相特征的早期融合以细化特征表示,克服双时相图像信息的孤立性;
Smart Images

Figure CN121837959B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of remote sensing image change detection and deep learning, specifically relating to an efficient remote sensing image change detection method based on a semantically guided hybrid network. Background Technology
[0002] Remote sensing change detection (RSCD) is one of the core technologies in the field of remote sensing, with broad application prospects in land cover analysis, disaster detection, urban planning, forest cover mapping, and resource management. The convenient acquisition of high-resolution optical remote sensing imagery has provided new opportunities for RSCD. The rich spatial details and clear texture information provided by high-resolution images greatly enhance the accuracy and semantic understanding capabilities of change detection. However, while providing rich ground feature information, noise and data complexity also increase dramatically. In the detailed interpretation of complex scenes, issues arise such as decreased image classification accuracy, the potential for information loss or misclassification due to shadows and occlusion, and insufficient understanding of the context by the model.
[0003] Deep learning is widely used in the RSCD field due to its powerful representation capabilities. The mainstream research architecture in this field is mainly based on two core pillars: Convolutional Neural Networks (CNN) and Transformers. Chinese Patent Publication No. CN119338780A, entitled "A Method and System for Remote Sensing Image Change Detection Based on Transformer and Graph Semantic Guidance," discloses a method and system for remote sensing image change detection based on Transformer and graph semantic guidance in the field of image processing technology. Its shortcomings include the difficulty of effectively capturing long-distance dependencies between pixels in remote sensing image change detection, and the tendency of direct subtraction in feature fusion to destroy feature structures and generate noise interference. As can be seen from the above, current algorithms using CNN or Transformer still suffer from high computational complexity and insufficient cross-scale and dual-temporal image information interaction. Common multi-scale fusion methods directly fuse shallow and deep features through splicing or addition. While simple, this method has poor adaptability and lacks a spatiotemporal interaction mechanism for dual-temporal images, limiting its ability to discriminate complex changing scenes.
[0004] This invention introduces an efficient cross-attention module to force feature interaction between bi-temporal images. Simultaneously, it sets a reduction ratio to lower the computational complexity of the transformer, integrating CNN into a multi-scale transformer encoder to construct an efficient multi-scale feature encoder. A strategy of fusing bi-temporal features by introducing deep features to guide shallow features and stacking pointwise convolutions is employed. Furthermore, joint loss and auxiliary loss are introduced to achieve high-accuracy change detection. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides an efficient remote sensing image change detection method based on a semantically guided hybrid network. It utilizes a hybrid architecture of CNN and Transformer to efficiently acquire multi-scale feature information, designs a semantically guided fusion module for cross-scale features, and uses stacked pointwise convolutions to fuse dual-temporal features, thereby significantly improving the accuracy of remote sensing image change detection.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0007] An efficient remote sensing image change detection method based on a semantically guided hybrid network, the method comprising the following steps:
[0008] Step S1: Input the dual-temporal remote sensing images into the high-efficiency cross-attention module. The dual-temporal remote sensing images include remote sensing image temporal phase 1 and remote sensing image temporal phase 2. This module forces the dual-temporal images to interact and compare, and sets the reduction ratio. To reduce the sequence length of keys and values and decrease the computational complexity of feature extraction, an efficient attention module is concatenated with random discarding to form a feature extraction module. Repeating this concatenation of the three modules generates feature information at three different scales.
[0009] Size is The characteristics, including phase 1 With phase 2 ;
[0010] Size is The characteristics, including phase 1 With phase 2 ;
[0011] Size is The characteristics, including phase 1 With phase 2 ,
[0012] in Indicates the image height. Indicates the image width;
[0013] Step S2 involves employing different fusion strategies for the extracted features at different scales, utilizing deep semantics to guide shallow features and leveraging deep, high-semantic features. , Guide and enhance mid-layer features , To obtain primary semantic enhancement features , Mid-layer features , Guided fusion of shallow high-resolution features , Obtain semantically guided fusion features , ;
[0014] Step S3, extract low-resolution dual-temporal features , Semantic guidance fusion features , and , These three different scales of features are used as inputs to stacked pointwise convolutional modules. For each scale of feature, the feature maps of corresponding time phase 1 and time phase 2 are concatenated along the channel dimension and input into a convolutional module consisting of two stacked pointwise convolutions. This enables the interaction of dual-temporal information and upsampling operations. The resulting features are then subjected to a difference operation to obtain the upsampled deep dual-temporal difference features. Mid-layer dual-temporal difference characteristics and shallow two-phase difference characteristics , Enhanced differential features are obtained through concatenated channel attention and spatial attention processing. ;
[0015] Step S4: Calculate using a fixed Sobel convolution kernel. , The gradients in the horizontal and vertical directions are then enhanced using depthwise dilated convolutions with kernel sizes of 3, 5, and 7 and dilation rates of 1, 2, and 3. , The weight of detail information during feature fusion, and the processed features and Fusion ;
[0016] Step S5: Construct a joint loss function using cross-entropy and the Dice function, and then use convolution to... , and Fusion ,Will The weighted graph generated by the prediction head is used as the main loss, and the following is utilized: The generated prediction weight map is used to construct the auxiliary loss.
[0017] Furthermore, in step S1, the remote sensing image is encoded into feature structures at three different scales, specifically as follows:
[0018] Dual-branch coding is used for dual-temporal remote sensing images, two-dimensional images , The image is segmented into fixed-size tiles using block embedding, where The number of feature channels is represented by the feature number, which is obtained after layer normalization. and ,in It refers to the number of tokens, where each token represents a feature unit corresponding to a segment of the image. As The source, setting the reduction ratio Using a step size of Downsampling is performed on the convolutional layers to obtain the dimensionality-reduced features. ,Will As and The source, the generation , and They are respectively:
[0019]
[0020]
[0021]
[0022] in, Indicates query characteristics, Indicates key features, Indicates value characteristics, express The projection weight matrix, express The projection weight matrix express The projection weight matrix, This indicates a normalization operation. This represents a 2D convolution operation. Indicates will The sequence is reshaped into Two-dimensional image;
[0023] calculate and The dot product of the two numbers is used to obtain the cross-attention weight matrix by applying Softmax normalization.
[0024]
[0025] in, Describes the dimension of K. express Matrix transpose;
[0026] The size is obtained by weighting V with attention weights, inserting a depthwise convolution between two fully connected layers of the MLP, and then concatenating a random dropout module. Dual-phase output characteristics and Repeat the above operation twice to obtain a size of Features , With size Features , , where MLP represents a feedforward neural network.
[0027] Furthermore, inserting a depthwise convolution between the two fully connected layers of the MLP specifically involves:
[0028] Used to introduce local spatial context information and assist the Transformer encoder module in obtaining feature information, the structure of the MLP is as follows:
[0029]
[0030]
[0031] in, This represents a 3×3 depthwise separable convolution operation. This indicates a feature flattening operation. This represents the intermediate feature connecting the attention module and the MLP. Represents the Gaussian error linear unit activation function. This represents the output characteristics of the efficient cross-attention module. The weight matrix represents the dimensionality reshaping. The bias term representing dimensional reshaping. Indicates the weights of the fully connected layer. This represents the bias term of the fully connected layer;
[0032] Furthermore, this MLP unit includes residual connections and introduces DropPath regularization and LayerNorm layer normalization.
[0033] Furthermore, the specific steps of step S2, deep semantic-guided multi-scale fusion, are as follows:
[0034] Low-resolution features are used to generate attention maps via a small convolutional network. and This attention map guides the retention of useful details in high-resolution features while removing useless shallow noise. and The features are upsampled and aligned with the spatial resolution of the previous scale of the input features, fused using residual connections, and then smoothed through convolutional layers to obtain semantically guided fused features. and .
[0035] Furthermore, the attention map is specifically represented as follows:
[0036] Attention Map and Represented as:
[0037]
[0038]
[0039] in, To represent the features of two images, This indicates that the activation function is used to generate an attention map that can be used as feature weights. This indicates that the small convolutional layer performs feature transformations to generate the attention map. Represents the linear rectified activation function. and This indicates that the features of the two images are the output of the low-resolution features processed by a small convolutional network, which is used to selectively activate the high-resolution features.
[0040] Furthermore, the specific steps of step S3, stacking point-by-point convolution, are as follows:
[0041] The three bi-temporal features at different scales are first spliced along the channel dimension to aggregate the features. :
[0042]
[0043] in, This indicates a channel-level concatenation operation. This represents the intermediate processing features of phase 1 of the dual-temporal remote sensing image. This represents the intermediate processing features of phase 2 of the dual-temporal remote sensing image;
[0044] Aggregation features Interactive features are obtained through two nonlinear transformations consisting of pointwise convolutions. :
[0045]
[0046]
[0047] in, This indicates a batch normalization operation. This represents a 1×1 convolution operation. Indicates intermediate interaction features;
[0048] The features are decoupled and upsampled, and then a difference operation is performed to obtain the upsampled deep dual-temporal difference features. Mid-layer dual-temporal difference characteristics Shallow two-phase difference characteristics , Enhanced differential features are obtained through concatenated channels and spatial attention processing. .
[0049] Furthermore, the upsampled deep dual-temporal difference features Mid-layer dual-temporal difference characteristics Shallow two-phase difference characteristics and enhanced differential features Specifically, it is expressed as follows:
[0050] Upsampled dual-temporal difference features , , Specifically, it is expressed as follows:
[0051]
[0052]
[0053] in, Representing the difference feature, This represents the absolute value operation. This indicates a convolution operation with a kernel size of 3, which aggregates features. Represented as:
[0054]
[0055] Difference features after channel spatial attention enhancement Specifically, it is expressed as follows:
[0056]
[0057] The calculation function of channel attention And spatial attention computation function Represented as:
[0058]
[0059]
[0060] in, Global average pooling operation, Global max pooling operation, Global mean pooling operation, This indicates a convolution operation with a kernel size of 7.
[0061] Furthermore, the specific steps of step S4, which utilizes the Sobel operator to enhance the details information, are as follows:
[0062] Three parallel depth-dilated convolution pairs are used. and Features were extracted from different receptive fields using convolutional kernel sizes of 3, 5, and 7, with corresponding dilation rates of 1, 2, and 3. The specific formula is as follows:
[0063]
[0064] in, This represents the intermediate features after depthwise separable convolution processing. This represents a depthwise separable convolution operation;
[0065] Calculate using a fixed Sobel convolution kernel respectively and The gradients in the horizontal and vertical directions are summed by taking their absolute values:
[0066]
[0067] in, This represents the detailed enhanced intermediate features obtained after processing by the Sobel operator. This indicates that a convolution operation is performed in the horizontal direction. This indicates that a convolution operation is performed in the vertical direction. Intermediate processing features representing the temporal-phase 2-fusion branch scale of dual-temporal remote sensing images;
[0068] The processed features are then fused, and a detail attention map is generated using the Sigmoid function. Detailed information is obtained by using residual joins. Detailed information of dual-temporal images The result obtained in step S3 Fusion, and the resulting features .
[0069] Furthermore, the specific formula for fusing the processed features is as follows:
[0070] Detail Attention Map Represented as:
[0071]
[0072] Detailed information is obtained by using residual joins. Represented as:
[0073]
[0074] The detailed information of the dual-temporal images and the information obtained in step S3 Fusion, and the resulting features Represented as:
[0075] .
[0076] Furthermore, the multi-scale auxiliary loss in step S5 is:
[0077] Using convolution , and Fusion :
[0078]
[0079]
[0080] in, This represents the fused features obtained after a full process including dual-temporal feature fusion, convolutional thinning, attention weighting, and detail enhancement.
[0081] and A weight map is generated using the prediction head. and The probability distribution diagram of actual changes We construct the main loss and auxiliary loss separately. The auxiliary loss provides additional and stronger gradient flow for the shallow and middle layers of the training network, and the constraints applied at different depths of the network have a similar effect to regularization. The overall loss function formula is:
[0082]
[0083] in, This represents the joint loss function consisting of binary cross-entropy (BCE) and Dice loss. The weighting coefficients for the auxiliary losses are represented by the following formula for the joint loss function:
[0084]
[0085] in, This represents the weighting parameter used to balance the proportions of the joint loss function;
[0086] Binary cross-entropy loss Represented as:
[0087]
[0088] in, H represents the total resolution of the image, and H and W represent the height and width. and They refer to spatial locations as By combining the ground truth values and the model's predicted output, the Dice loss function is introduced to compensate for the susceptibility of the BCE loss to background interference. The specific formula is as follows:
[0089]
[0090] Where N is the total number of pixels in the ground view. and These represent the values of the nth pixel in the predicted map and the ground reality label, respectively.
[0091] Compared with existing technologies, it has the following technical advantages:
[0092] 1) This invention proposes an efficient remote sensing image change detection method based on a semantically guided hybrid network. Addressing the high computational cost of existing hybrid architectures, this method reduces computational overhead during the encoding stage. and To improve spatial resolution, minimize computational overhead, utilize cross-attention, and force early fusion of bi-temporal features to refine feature representations, thereby overcoming the isolation of bi-temporal image information;
[0093] 2) The present invention proposes an efficient remote sensing image change detection method based on a semantically guided hybrid network. This method uses the semantic information of deep features to guide shallow features, focusing on the information of real changes, and effectively solves the problem of introducing shallow noise.
[0094] 3) The present invention proposes an efficient remote sensing image change detection method based on semantically guided hybrid network. By fusing dual-temporal features of the same scale through stacked pointwise convolution, features at different levels are mapped to the same latent space, reducing semantic differences and enhancing the nonlinear expressive power of the model.
[0095] 4) The present invention proposes an efficient remote sensing image change detection method based on semantically guided hybrid networks, which introduces multi-scale deep dilated convolution and fixed Sobel convolution kernels to enhance the feature information of small objects that are easily ignored by the hybrid architecture.
[0096] 5) The present invention proposes an efficient remote sensing image change detection method based on semantically guided hybrid network. It adopts a joint loss composed of BCE and Dice to effectively alleviate the class imbalance problem, introduces auxiliary loss to accelerate the overall convergence speed of the model, and effectively prevents the model from overfitting. Attached Figure Description
[0097] Figure 1 This is a flowchart of the steps of the present invention;
[0098] Figure 2 This is a diagram of the overall framework of the ESG-Net model of this invention;
[0099] Figure 3 This is a schematic diagram of the efficient cross-attention framework of the present invention;
[0100] Figure 4 This is a schematic diagram of semantic guidance and stacked pointwise convolution of the present invention;
[0101] Figure 5 This is a detailed feature diagram of the present invention. Detailed Implementation
[0102] The present invention is described below based on embodiments, but the invention is not limited to these embodiments. In the detailed description of the invention below, certain specific details are described in detail. Those skilled in the art can fully understand the invention even without these detailed descriptions. To avoid obscuring the essence of the invention, well-known methods, processes, flows, elements, and circuits are not described in detail. To make the objectives, technical solutions, and advantages of the invention clearer, the invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention.
[0103] This embodiment relates to an efficient remote sensing image change detection method based on a semantically guided hybrid network. It acquires multi-scale feature information by constructing an efficient cross-attention mechanism. Through semantic guidance of deep features and stacked pointwise convolution operations, it breaks down information silos between images of different scales and in two temporal phases. Furthermore, it introduces joint loss and auxiliary loss to effectively alleviate class imbalance and overfitting, thus mitigating the high computational complexity and insufficient multi-scale fusion issues of hybrid architectures. Figure 1 The diagram shown is a flowchart of the steps of the present invention, which includes the following steps:
[0104] Step S1: Input the dual-temporal remote sensing images into the high-efficiency cross-attention module. The dual-temporal remote sensing images include remote sensing image temporal phase 1 and remote sensing image temporal phase 2. This module forces the dual-temporal images to interact and compare, and sets the reduction ratio. To reduce the sequence length of keys and values and decrease the computational complexity of feature extraction, an efficient attention module is concatenated with a random dropout module to form a single feature extraction module. Repeating this concatenation of the three modules generates feature information at three different scales.
[0105] Generate feature information at three different scales:
[0106] Size is The characteristics, including phase 1 With phase 2 ;
[0107] Size is The characteristics, including phase 1 With phase 2 ;
[0108] Size is The characteristics, including phase 1 With phase 2 ,
[0109] in Indicates the image height. Indicates the image width;
[0110] Step S2 involves employing different fusion strategies for the extracted features at different scales, utilizing deep semantics to guide shallow features and leveraging deep, high-semantic features. , Guide and enhance mid-level features , To obtain primary semantic enhancement features , Mid-layer features , Guided fusion of shallow high-resolution features , Obtain semantically guided fusion features , ;
[0111] Step S3, extract low-resolution dual-temporal features , Semantic guidance fusion features , and , These three different scales of features are used as inputs to stacked pointwise convolutional modules. For each scale of feature, the feature maps of corresponding time phase 1 and time phase 2 are concatenated along the channel dimension and input into a convolutional module consisting of two stacked pointwise convolutions. This enables the interaction of dual-temporal information and upsampling operations. The resulting features are then subjected to a difference operation to obtain the upsampled deep dual-temporal difference features. Mid-layer dual-temporal difference characteristics Difference characteristics between shallow and two phases , Enhanced differential features are obtained through concatenated channel attention and spatial attention processing. ;
[0112] Step S4: Calculate using a fixed Sobel convolution kernel. , The gradients in the horizontal and vertical directions are then enhanced using depthwise dilated convolutions with kernel sizes of 3, 5, and 7 and dilation rates of 1, 2, and 3. , The weight of detail information during feature fusion, and the processed features and Fusion ;
[0113] Step S5: Construct a joint loss function using cross-entropy and the Dice function, and then use convolution to... , and Fusion ,Will The weighted graph generated by the prediction head is used as the main loss, and the following is utilized: The generated prediction weight map is used to construct the auxiliary loss.
[0114] Furthermore, in step S1, the remote sensing image is encoded into feature structures at three different scales, specifically as follows:
[0115] For dual-temporal remote sensing images, dual-branch coding is used, such as Figure 2 As shown, a two-dimensional image , The image is segmented into fixed-size tiles using block embedding, where The number of feature channels is represented by the feature number, which is obtained after layer normalization. and ,in It refers to the number of tokens, where each token represents a feature unit corresponding to a segment of the image, such as... Figure 3 As shown, As The source, setting the reduction ratio Using a step size of Downsampling is performed on the convolutional layers to obtain the dimensionality-reduced features. ,Will As and The source, the generation , and They are respectively:
[0116]
[0117]
[0118]
[0119] in, Indicates query characteristics, Indicates key features, Indicates value characteristics, express The projection weight matrix, express The projection weight matrix express The projection weight matrix, This indicates a normalization operation. This represents a 2D convolution operation. Indicates will The sequence is reshaped into Two-dimensional image;
[0120] calculate and The dot product of the two numbers is used to obtain the cross-attention weight matrix by applying Softmax normalization.
[0121]
[0122] in, Describes the dimension of K. express Matrix transpose;
[0123] The size is obtained by weighting V with attention weights, inserting a depthwise convolution between two fully connected layers of the MLP, and then concatenating a random dropout module. Dual-phase output characteristics and Repeat the above operation twice to obtain a size of Features , With size Features , , where MLP represents a feedforward neural network.
[0124] Furthermore, inserting a depthwise convolution between the two fully connected layers of the MLP specifically involves:
[0125] Used to introduce local spatial context information and assist the Transformer encoder module in obtaining feature information, the structure of the MLP is as follows:
[0126]
[0127]
[0128] in, This indicates a 3×3 depth separable convolution operation. This indicates a feature flattening operation. This represents the intermediate feature connecting the attention module and the MLP. Represents the Gaussian error linear unit activation function. This represents the output characteristics of the efficient cross-attention module. The weight matrix represents the dimensionality reshaping. The bias term representing dimensional reshaping. Indicates the weights of the fully connected layer. This represents the bias term of the fully connected layer;
[0129] Furthermore, this MLP unit contains residual connections, such as Figure 2 As shown, DropPath regularization and LayerNorm layer normalization are introduced simultaneously.
[0130] Furthermore, the specific steps of step S2, deep semantic-guided multi-scale fusion, are as follows:
[0131] like Figure 4 As shown, low-resolution features are used to generate attention maps through a small convolutional network. and If a pixel value approaches 1, that information is retained; conversely, if it approaches 0, that feature is suppressed. (Attention map) and The features are upsampled and aligned with the spatial resolution of the previous scale of the input features, fused using residual connections, and then smoothed through convolutional layers to obtain semantically guided fused features. and .
[0132] Furthermore, the attention map is specifically represented as follows:
[0133] Attention Map and Represented as:
[0134]
[0135]
[0136] in, To represent the features of two images, This indicates that the activation function is used to generate an attention map that can be used as feature weights. This indicates that the small convolutional layer performs feature transformations to generate the attention map. Represents the linear rectified activation function. and This indicates that the features of the two images are the output of the low-resolution features processed by a small convolutional network, which is used to selectively activate the high-resolution features.
[0137] Furthermore, the specific steps of step S3, stacking point-by-point convolution, are as follows:
[0138] like Figure 4 As shown, by stitching together bi-temporal features of the same scale along the channel dimension, aggregated features are obtained. :
[0139]
[0140] in, This indicates a channel-level concatenation operation. This represents the intermediate processing features of phase 1 of the dual-temporal remote sensing image. This represents the intermediate processing features of phase 2 of the dual-temporal remote sensing image;
[0141] Aggregation features Interactive features are obtained through two nonlinear transformations consisting of pointwise convolutions. :
[0142]
[0143]
[0144] in, This indicates a batch normalization operation. This represents a 1×1 convolution operation. Indicates intermediate interaction features;
[0145] like Figure 2 The existence of channel and spatial attention is to further refine the multi-scale fused features and more efficiently capture the content and positional dependencies of the features. To avoid the loss of original feature information, residual connections are introduced to decouple and upsample the features, followed by a difference operation to obtain the upsampled deep dual-temporal difference features. Mid-layer dual-temporal difference characteristics Difference characteristics between shallow and two phases , The enhanced differential features are obtained through cascaded channel and spatial attention processing. .
[0146] Furthermore, the deep dual-temporal difference features after upsampling in step S3 Mid-layer dual-temporal difference characteristics Shallow two-phase difference characteristics and enhanced differential features Specifically, it is expressed as follows:
[0147] Upsampled dual-temporal difference features , , Specifically, it is expressed as follows:
[0148]
[0149]
[0150] in, Representing the difference feature, This represents the absolute value operation. This indicates a convolution operation with a kernel size of 3, which aggregates features. Represented as:
[0151]
[0152] Difference features after channel spatial attention enhancement Specifically, it is expressed as follows:
[0153]
[0154] The calculation function of channel attention And spatial attention computation function Represented as:
[0155]
[0156]
[0157] in, Global average pooling operation, Global max pooling operation, Global mean pooling operation, This indicates a convolution operation with a kernel size of 7. Channel attention uses average pooling and MLP to obtain content weights of the feature map, while spatial attention obtains position weight information of the feature map by concatenating the maximum and mean values and then passing them through a convolutional layer.
[0158] Furthermore, the specific steps of step S4, which utilizes the Sobel operator to enhance the details information, are as follows:
[0159] like Figure 5 As shown, three parallel depth-dilated convolution pairs are used. and Features were extracted from different receptive fields using convolutional kernel sizes of 3, 5, and 7, with corresponding dilation rates of 1, 2, and 3. The specific formula is as follows:
[0160]
[0161] in, This represents the intermediate features after depthwise separable convolution processing. This represents a depthwise separable convolution operation;
[0162] Calculate using a fixed Sobel convolution kernel respectively and The gradients in the horizontal and vertical directions are summed by taking their absolute values:
[0163]
[0164] in, This represents the detailed enhanced intermediate features obtained after processing by the Sobel operator. This indicates that a convolution operation is performed in the horizontal direction. This indicates that a convolution operation is performed in the vertical direction. Intermediate processing features representing the temporal-phase 2-fusion branch scale of dual-temporal remote sensing images;
[0165] The processed features are then fused, and a detail attention map is generated using the Sigmoid function. Detailed information is obtained by using residual joins. Detailed information of dual-temporal images The result obtained in step S3 Fusion, and the resulting features .
[0166] Furthermore, the specific formula for fusing the processed features is as follows:
[0167] Detail Attention Map Represented as:
[0168]
[0169] Detailed information is obtained by using residual joins. Represented as:
[0170]
[0171] like Figure 2 After processing by the Sobel operator enhancement detail information module, the detail information of the dual-temporal images is compared with that obtained in step S3. Fusion, and the resulting features Represented as:
[0172] .
[0173] Furthermore, the multi-scale auxiliary loss in step S5 is:
[0174] Introducing the joint loss consisting of BCE and Dice loss, and and A weight map is generated using the prediction head. and Construct a joint loss and use convolution to... , and Fusion :
[0175]
[0176]
[0177] in, This represents the fused features obtained after a full process including dual-temporal feature fusion, convolutional thinning, attention weighting, and detail enhancement.
[0178] and A weight map is generated using the prediction head. and The probability distribution diagram of actual changes We construct the main loss and auxiliary loss separately. The auxiliary loss provides additional and stronger gradient flow for the shallow and middle layers of the training network, and the constraints applied at different depths of the network have a similar effect to regularization. The overall loss function formula is:
[0179]
[0180] in, This represents the joint loss function consisting of binary cross-entropy (BCE) and Dice loss. The weighting coefficient for the auxiliary loss is set to 0.8 in this invention, which maximizes the model's performance. The specific formula for the joint loss function is as follows:
[0181]
[0182] in, This represents the weighting parameter used to balance the proportions of the joint loss function. In this invention, setting it to 1 yields the best results.
[0183] Binary cross-entropy loss Represented as:
[0184]
[0185] in, H represents the total resolution of the image, and H and W represent the height and width. and They refer to spatial locations as By combining the ground truth values and the model's predicted output, the Dice loss function is introduced to compensate for the susceptibility of the BCE loss to background interference. The specific formula is as follows:
[0186]
[0187] Where N is the total number of pixels in the ground view. and represents the value of the nth pixel in the predicted map and the ground reality label, respectively. The joint loss constructed using BCE and Dice alleviates the class imbalance that often occurs in change detection.
[0188] Compared to other RSCD methods using a hybrid CNN and Transformer architecture, this invention reduces computational complexity while addressing the issue of insufficient information exchange during multi-scale feature fusion. Furthermore, it mitigates the potential damage to detail information caused by downsampling operations in CNNs and patch processing in Transformers, effectively improving change detection accuracy. The technical means disclosed in this invention are not limited to those described in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered within the scope of protection of this invention.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A highly efficient remote sensing image change detection method based on a semantically guided hybrid network, characterized in that, The method includes the following steps: Step S1: Input the dual-temporal remote sensing images into the high-efficiency cross-attention module. The dual-temporal remote sensing images include remote sensing image temporal phase 1 and remote sensing image temporal phase 2. This module forces the dual-temporal remote sensing images to interact and compare, and sets the reduction ratio. To reduce the sequence length of keys and values and decrease the computational complexity of feature extraction, an efficient attention module is concatenated with random discarding to form a feature extraction module. Repeating this concatenation of the three modules generates feature information at three different scales. Size is The characteristics, including phase 1 With phase 2 ; Size is The characteristics, including phase 1 With phase 2 ; Size is The characteristics, including phase 1 With phase 2 , in Indicates the image height. Indicates the image width; Step S2 involves employing different fusion strategies for the extracted features at different scales, utilizing deep semantics to guide shallow features and leveraging deep, high-semantic features. , Guide and enhance mid-layer features , To obtain primary semantic enhancement features , Mid-layer features , Guided fusion of shallow high-resolution features , Obtain semantically guided fusion features , ; Step S3, extract low-resolution dual-temporal features , Semantic guidance fusion features , and , These three different scales of features are used as inputs to stacked pointwise convolutional modules. For each scale of feature, the feature maps of corresponding time phase 1 and time phase 2 are concatenated along the channel dimension and input into a convolutional module consisting of two stacked pointwise convolutions. This enables the interaction of dual-temporal information and upsampling operations. The resulting features are then subjected to a difference operation to obtain the upsampled deep dual-temporal difference features. Mid-layer dual-temporal difference characteristics and shallow two-phase difference characteristics , Enhanced differential features are obtained through concatenated channel attention and spatial attention processing. ; Step S4: Calculate using a fixed Sobel convolution kernel. , The gradients in the horizontal and vertical directions are then enhanced using depthwise dilated convolutions with kernel sizes of 3, 5, and 7 and dilation rates of 1, 2, and 3. , The weight of detail information during feature fusion, and the processed features and Fusion ; Step S5: Construct a joint loss function using cross-entropy and the Dice function, and then use convolution to... , and Fusion ,Will The weighted graph generated by the prediction head is used as the main loss, and the following is utilized: The generated prediction weight map is used to construct the auxiliary loss.
2. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 1, characterized in that, In step S1, the remote sensing image is encoded into feature structures at three different scales, specifically as follows: Dual-branch coding is used for dual-temporal remote sensing images, two-dimensional images , The image is segmented into fixed-size tiles using block embedding, where The number of feature channels is represented by the feature number, which is obtained after layer normalization. and ,in It refers to the number of tokens, where each token represents a feature unit corresponding to a segment of the image. As The source, setting the reduction ratio Using a step size of Downsampling is performed on the convolutional layers to obtain the dimensionality-reduced features. ,Will As and The source, the generation , and They are respectively: in, Indicates query characteristics, Indicates key features, Indicates value characteristics, express The projection weight matrix, express The projection weight matrix express The projection weight matrix, This indicates a normalization operation. This represents a 2D convolution operation. Indicates will The sequence is reshaped into Two-dimensional image; calculate and The dot product of the two numbers is used to obtain the cross-attention weight matrix by applying Softmax normalization. in, Describes the dimension of K. express Matrix transpose; The size is obtained by weighting V with attention weights, inserting a depthwise convolution between two fully connected layers of the MLP, and then concatenating a random dropout module. Dual-phase output characteristics and Repeat the above operation twice to obtain a size of Features , With size Features , , where MLP represents a feedforward neural network.
3. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 2, characterized in that, The insertion of a depthwise convolution between the two fully connected layers of the MLP is specifically as follows: Used to introduce local spatial context information and assist the Transformer encoder module in obtaining feature information, the structure of the MLP is as follows: in, This represents a 3×3 depthwise separable convolution operation. This indicates a feature flattening operation. This represents the intermediate feature connecting the attention module and the MLP. Represents the Gaussian error linear unit activation function. This represents the output characteristics of the efficient cross-attention module. The weight matrix represents the dimensionality reshaping. The bias term representing dimensional reshaping. Indicates the weights of the fully connected layer. This represents the bias term of the fully connected layer; Furthermore, this MLP unit includes residual connections and introduces DropPath regularization and LayerNorm layer normalization.
4. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 1, characterized in that, The specific steps of step S2, deep semantic-guided multi-scale fusion, are as follows: Low-resolution features are used to generate attention maps via a small convolutional network. and This attention map guides the retention of useful details in high-resolution features while removing useless shallow noise. and The features are upsampled and aligned with the spatial resolution of the previous scale of the input features, fused using residual connections, and then smoothed through convolutional layers to obtain semantically guided fused features. and .
5. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 4, characterized in that, The attention map is specifically represented as follows: Attention Map and Represented as: in, To represent the features of two images, This indicates that the activation function is used to generate an attention map that can be used as feature weights. This indicates that the small convolutional layer performs feature transformations to generate the attention map. Represents the linear rectified activation function. and This indicates that the features of the two images are the output of the low-resolution features processed by a small convolutional network, which is used to selectively activate the high-resolution features.
6. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 1, characterized in that, The specific steps of step S3, stacked point-by-point convolution, are as follows: The three bi-temporal features at different scales are first spliced along the channel dimension to aggregate the features. : in, This indicates a channel-level concatenation operation. This represents the intermediate processing features of phase 1 of the dual-temporal remote sensing image. This represents the intermediate processing features of phase 2 of the dual-temporal remote sensing image; Aggregation features Interactive features are obtained through two nonlinear transformations consisting of pointwise convolutions. : in, This indicates a batch normalization operation. This represents a 1×1 convolution operation. Indicates intermediate interaction features; The features are decoupled and upsampled, and then a difference operation is performed to obtain the upsampled deep dual-temporal difference features. Mid-layer dual-temporal difference characteristics Shallow two-phase difference characteristics , Enhanced differential features are obtained through concatenated channels and spatial attention processing. .
7. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 6, characterized in that, The upsampled deep bi-temporal difference features Mid-layer dual-temporal difference characteristics Shallow two-phase difference characteristics and enhanced differential features Specifically, it is expressed as follows: Upsampled dual-temporal difference features , , Specifically, it is expressed as follows: in, Representing the difference feature, This represents the absolute value operation. This indicates a convolution operation with a kernel size of 3, which aggregates features. Represented as: Difference features after channel spatial attention enhancement Specifically, it is expressed as follows: The calculation function of channel attention And spatial attention computation function Represented as: in, Global average pooling operation, Global max pooling operation, Global mean pooling operation, This indicates a convolution operation with a kernel size of 7.
8. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 1, characterized in that, The specific steps of step S4, which utilizes the Sobel operator to enhance the detail information, are as follows: Three parallel depth-dilated convolution pairs are used. and Features were extracted from different receptive fields using convolutional kernel sizes of 3, 5, and 7, with corresponding dilation rates of 1, 2, and 3. The specific formula is as follows: in, This represents the intermediate features after depthwise separable convolution processing. This represents a depthwise separable convolution operation; Calculate using a fixed Sobel convolution kernel respectively and The gradients in the horizontal and vertical directions are summed by taking their absolute values: in, This represents the detailed enhanced intermediate features obtained after processing by the Sobel operator. This indicates that a convolution operation is performed in the horizontal direction. This indicates that a convolution operation is performed in the vertical direction. Intermediate processing features representing the temporal-phase 2-fusion branch scale of dual-temporal remote sensing images; The processed features are then fused, and a detail attention map is generated using the Sigmoid function. Detailed information is obtained by using residual joins. Detailed information of dual-temporal images The result obtained in step S3 Fusion, and the resulting features .
9. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 8, characterized in that, The specific formula for fusing the processed features is as follows: Detail Attention Map Represented as: Detailed information is obtained by using residual joins. Represented as: The detailed information of the dual-temporal images and the information obtained in step S3 Fusion, and the resulting features Represented as: 。 10. The efficient remote sensing image change detection method based on a semantically guided hybrid network according to claim 1, characterized in that, The multi-scale auxiliary loss in step S5 is: Using convolution , and Fusion : in, This represents the fused features obtained after a full-process processing including dual-temporal feature fusion, convolutional refinement, attention weighting, and detail enhancement. and A weight map is generated using the prediction head. and The probability distribution diagram of actual changes The main loss and auxiliary loss are constructed separately, and the overall loss function formula is as follows: in, This represents the joint loss function consisting of binary cross-entropy (BCE) and Dice loss. The weighting coefficients for the auxiliary losses are represented by the following formula for the joint loss function: in, This represents the weighting parameter used to balance the proportions of the joint loss function; Binary cross-entropy loss Represented as: in, H represents the total resolution of the image, and H and W represent the height and width. and They refer to spatial locations as By combining the ground truth values and the model's predicted output, the Dice loss function is introduced to compensate for the susceptibility of the BCE loss to background interference. The specific formula is as follows: Where N is the total number of pixels in the ground view. and These represent the values of the nth pixel in the predicted map and the ground reality label, respectively.
Citation Information
Patent Citations
Remote sensing image change detection method and system based on Transform and graph semantic guidance
CN119338780A
Remote sensing image change detection system and method based on multi-modal deep learning
CN120147895A
Method for classifying hyperspectral images on basis of adaptive multi-scale feature extraction model
US20230252761A1