A method for intelligent interpretation of dual-temporal remote sensing images
Patent Information
- Application Number
- CN202610130644.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-01-30
AI Technical Summary
[0006]基于此,本发明提供了一种双时相遥感影像智能解译方法,以解决背景技术中所提到的复杂场景下变化目标尺度跨度大、语义差异微弱以及边界不规则等导致的变化检测不准确和变化目标边界不完整的问题
[0076] As can be seen from the above technical solution, the beneficial effects of the intelligent interpretation method for dual-temporal remote sensing images proposed in this invention are as follows:
Smart Images

Figure CN121982453B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent interpretation methods for remote sensing images, and particularly to an intelligent interpretation method for dual-temporal remote sensing images. Background Technology
[0002] Intelligent interpretation methods for remote sensing images aim to accurately identify and locate changes in land cover and human activities by analyzing remote sensing images of the same geographic area acquired at different times. As a key technology in Earth observation and environmental monitoring, intelligent interpretation methods for remote sensing images play an irreplaceable role in practical applications such as urban expansion monitoring, land use and cover change analysis, disaster assessment, ecological environmental protection, and land resource management. With the continuous improvement of the spatial resolution of remote sensing images, ground features exhibit richer textures and geometric characteristics, especially in densely built-up urban areas, farmland areas, and complex natural landscapes. Interpretation targets often exhibit characteristics such as varying scales, irregular shapes, and subtle semantic differences. How to accurately extract change information under complex background interference while maintaining the geometric integrity of the changed target boundaries has become a core challenge of common concern in both academia and engineering applications.
[0003] With the rapid development of deep learning technology, intelligent interpretation methods for remote sensing images based on convolutional neural networks (CNNs) and Transformers have become the mainstream research paradigm. These deep learning methods, with their powerful feature extraction and context modeling capabilities, have effectively improved change detection accuracy. However, there is still room for improvement in perceiving fine-grained changes in complex scenes and in finely characterizing change boundaries.
[0004] Visual foundation models (VBMs) can effectively model complex ground structures and high-level semantic information, providing a novel research paradigm for intelligent interpretation of remote sensing images. VBMs refer to general visual representation models pre-trained on large-scale general image or multimodal datasets through self-supervised, weakly supervised, or contrastive learning methods. Typical examples include SAM, CLIP, and DINOv2. Benefiting from large-scale data-driven pre-training, VBMs can learn visual representations with strong generalization capabilities and cross-task transfer potential. However, directly applying VBMs to change detection tasks in intelligent remote sensing image interpretation still has significant limitations: on the one hand, the high-level features output by VBMs typically focus on semantic consistency and global structure modeling, with limited ability to represent local details. On the other hand, directly transferring general VBMs often ignores the stringent requirements of change detection for spatial positioning accuracy and boundary geometric integrity, easily leading to problems such as blurred boundaries, missing details, or overly smoothed results. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] Based on this, the present invention provides an intelligent interpretation method for dual-temporal remote sensing images to solve the problems of inaccurate change detection and incomplete boundary of changed targets caused by large scale span of changing targets, weak semantic differences and irregular boundaries in complex scenes mentioned in the background art.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the present invention provides an intelligent interpretation method for dual-temporal remote sensing images, comprising:
[0009] Step S1: Construct a boundary-constrained remote sensing image change detection model based on a visual fundamental model;
[0010] Step S2: Based on the remote sensing image dataset, train the change detection model to obtain a trained change detection model;
[0011] Step S3: Input remote sensing images from different time phases into the trained change detection model and obtain the interpretation results;
[0012] The boundary-constrained remote sensing image change detection model BCnet based on the visual basic model includes: a SAM encoder module, a difference detail enhancement module, a multi-scale edge enhancement module, an edge feature aggregation module, and a change detection head; wherein, the edge feature aggregation module includes an edge detection head; the change detection head includes a convolution with a 3×3 kernel.
[0013] The dual-temporal remote sensing images are input into the SAM encoder module to extract shallow semantic features and four levels of semantic features; the four levels of semantic features are input into the difference detail enhancement module to obtain four difference features; the shallow semantic features and the four difference features are input into the multi-scale edge enhancement module to obtain multi-scale edge features; on one hand, the multi-scale edge features are input into the edge detection head to obtain the edge mask; on the other hand, the multi-scale edge features are input into the edge feature aggregation module to obtain the edge aggregation features, and the edge aggregation features are input into the change detection head to obtain the change mask.
[0014] In the SAM encoder module, the single-input SAM image encoder is improved into a twin structure with two dual-temporal remote sensing image inputs. That is, the dual-temporal remote sensing images T1 and T2 are respectively input into the first SAM image encoder and the second SAM image encoder, and the two SAM image encoders share weights.
[0015] The difference detail enhancement module comprises four difference detail enhancement units (DDEUs). These four units, namely the first, second, third, and fourth DDEUs, are connected to the 6th, 12th, 18th, and 24th Transformer layers in the SAM image encoder, respectively, which have the global receptive field. Specifically, the remote sensing image T1 is input into the first SAM image encoder, and the difference detail enhancement units (DDEUs) are connected to the 6th, 12th, 18th, and 24th Transformer layers in its global receptive field. Semantic features are obtained from each Transformer layer. The remote sensing image T2 is input into the second SAM image encoder, from its corresponding... Semantic features are obtained from a Transformer layer. ;in, ;
[0016] Subsequently, feature pairs and The first input DDEUs and generate output. ;in, Indicates the level number of the Transformer. Indicates the cell number of the DDEU, and The specific correspondence is as follows: [The rest of the text appears to be a list of characters and symbols, possibly related to a specific project or event.] and As the input of the first DDEU, the output ;Will and As the input to the second DDEU, the output ;Will and As the input of the third DDEU, the output ;Will and As the input of the fourth DDEU, the output .
[0017] Furthermore, when describing the structure of DDEU, the following is used: and Representing the two input features of DDEU, using To represent its output features, DDEU contains two parallel branches: a detail guidance branch and a difference modeling branch;
[0018] In the detailed guidance branch, the input features... and The splicing is performed along the channel dimension to obtain the spliced features. ,in, ;Right now:
[0019]
[0020] in, Indicates splicing;
[0021] Subsequently, the spliced features are processed through two convolutional operations. Enhancement and compression are performed: the first convolutional layer is used to enhance features, and the second convolutional layer is used to reduce the number of channels to a smaller dimension. , obtain detailed guided features ;
[0022] In the differential modeling branch, the element-wise absolute difference between the two input features is first calculated to generate differential features. :
[0023]
[0024] To simultaneously capture the distribution characteristics of change information in both spatial and channel dimensions, The data is fed into two parallel sub-paths: on the channel sub-path, channel statistical features are obtained by global max pooling along the spatial axis. Spatial response features are obtained by compressing along the channel axis using global max pooling on the spatial sub-path. ;
[0025] Subsequently, channel attention weights and spatial attention weights are generated through convolution operations and a sigmoid activation function, respectively. and :
[0026]
[0027] in, This represents the Sigmoid function, used to constrain weights within the [0,1] interval;
[0028] Finally, the aforementioned attention weights are applied to the detail-guided features respectively. Enhanced differential features are generated through weighted fusion. :
[0029] .
[0030] Furthermore, the input to the multi-scale edge enhancement module MEEM includes semantic fusion features from the first Transformer layer in the SAM image encoder. and four difference features from the output of the difference detail enhancement module. , , , In this process, the remote sensing image T1 is input into the first SAM image encoder, which extracts shallow semantic features from the first Transformer layer, which has a global receptive field. The remote sensing image T2 is input into the second SAM image encoder to obtain shallow semantic features from its corresponding first Transformer layer. ; and Element-wise addition ;
[0031] The MEEM processing flow is as follows:
[0032] (1) Extraction of prior features of shallow edges:
[0033] MEEM first analyzes shallow semantic fusion features. Multi-scale convolution processing is performed: a set of parallel 3×3 convolutions combined with upsampling are used to convert them into shallow edge prior features with different spatial resolutions but uniform channel dimensions. Specifically, The first sub-path is obtained by using four parallel sub-paths, each of which undergoes an 8x upsampling and a 3×3 convolution. The second sub-path is obtained by 4x upsampling and 3×3 convolution. The third sub-path is obtained by 2x upsampling and 3×3 convolution. The fourth sub-path is obtained by upsampling by 1x and convolution by 3×3. ;
[0034] (2) Alignment and enhancement of differential features:
[0035] Difference features from the output of the difference detail enhancement module First, transpose convolution is used to enhance its scale information representation capability. Then, 1×1 convolution is used to compress and align the channel dimensions to obtain a multi-scale differential feature representation with a uniform number of channels. :
[0036]
[0037] in, Indicates transposed convolution; Represents convolution;
[0038] (3) Feature fusion
[0039] shallow edge prior features Multiscale differential features By adding dimensions, we obtain the fused features. :
[0040]
[0041] (4) Two-way interaction
[0042] A two-way interactive strategy combining top-down and bottom-up approaches is adopted: high-level semantic features are downsampled and passed to the low-level layer to suppress background noise; low-level detail features are upsampled to supplement high-frequency edge information.
[0043] Specifically , By downsampling and Adding them together gives , By downsampling and Adding them together gives , By downsampling and Adding them together gives ; , Through upsampling and Adding them together gives , Through upsampling and Adding them together gives , Through upsampling and Adding them together gives ;
[0044] (5) Edge enhancement and multi-scale feature integration
[0045] After each level of fusion in the bidirectional interaction, a multilayer perceptron is introduced to perform nonlinear mapping and channel recalibration of the features to further enhance the boundary response and obtain edge-enhanced features at each scale. ,Right now Obtained through a multilayer perceptron ;
[0046] Edge enhancement features output at each scale The resolution is uniformly adjusted to 128×128, and the data is stitched together along the channel dimension to form the final multi-scale edge feature representation:
[0047] .
[0048] Furthermore, the edge feature aggregation module EFAM uses the output features of MEEM. As input, EFAM includes an edge detection head and two convolutional layers; The path taken by the edge detection head is called the edge detection head path. The path that directly passes through two convolutional layers is called the original feature path;
[0049] In the edge detection head, the input features Max pooling is performed along both the horizontal and vertical directions to obtain statistical features sensitive to both directions. and This operation not only captures long-range contextual dependencies in the horizontal direction, but also preserves positional relationships in the vertical direction, which helps the model to more accurately locate the spatial extent of changing objects.
[0050] Then, and The data are fed into convolutional layers, and normalized spatial horizontal attention weights are generated using the Sigmoid activation function. Attention weights in the perpendicular direction of space ;
[0051]
[0052] in, This represents the Sigmoid activation function; Represents convolution;
[0053] After obtaining the attention weights, EFAM adaptively reweights the input features through element-wise multiplication to achieve refined focusing on regions of change; as shown below:
[0054]
[0055] To further enhance the consistency and discriminative power of local features, EFAM reweights the features. Apply a convolutional transformation;
[0056] Meanwhile, in the original feature path, double convolution pairs are used. Nonlinear projection is performed to ensure semantic alignment between the original feature path and the edge detection head path. To prevent excessive boundary guidance from damaging the original multi-scale edge features, the module employs a residual connection mechanism to fuse the edge detection head path and the original feature path to obtain the edge aggregation feature. As shown below:
[0057] .
[0058] Furthermore, an edge feature constraint strategy EFCS was designed, which includes three parts: edge label generation, edge supervision loss, and edge back-guidance.
[0059] (1) Edge label generation
[0060] A morphology-based automatic edge label generation strategy is designed, which constructs reliable boundary supervision signals from binary ground truth labels using an edge generator. Specifically, firstly, Gaussian filtering is applied to the ground truth labels of changing regions to smooth them, suppressing noise and mitigating the discreteness of labeled edges. Subsequently, the initial boundaries of changing regions are extracted from the smoothed ground truth labels using the Canny edge detection operator. To enhance the continuity and fault tolerance of the boundaries, an expansion operation is performed on the initial boundaries to obtain the edge labels. ;
[0061] (2) Edge supervision loss
[0062] During training, BCnet incorporates multi-scale edge features. Input edge detection head branch, output boundary prediction probability map This refers to the edge mask; to balance pixel-level classification accuracy with the overall shape's structural consistency, an edge loss function combining binary cross-entropy loss (BCE Loss) and Dice loss is used. :
[0063]
[0064] in, This represents the binary cross-entropy loss; Indicates Dice loss;
[0065] (3) Edge reverse guidance
[0066] Edge aggregation features The change prediction result, i.e., the change mask, is jointly input into the change detection head branch along with the boundary prediction probability map, thus providing guidance at the feature level. The change prediction result output by the change detection head branch... Binary cross-entropy loss is used to supervise the change region; as shown below:
[0067]
[0068] in, This represents the true result of the change, i.e., the true label;
[0069] Combining the change detection task and the boundary constraint task, the model is trained using the following joint loss function:
[0070]
[0071] in, This represents a tradeoff coefficient used to balance the accuracy of change detection with the strength of boundary constraints.
[0072] Furthermore, It is 0.3.
[0073] Furthermore, a LoRA adapter is introduced to fine-tune the parameters of the SAM image encoder.
[0074] Furthermore, the remote sensing image dataset in step S2 is LEVIR-CD.
[0075] (III) Beneficial Effects
[0076] As can be seen from the above technical solution, the beneficial effects of the intelligent interpretation method for dual-temporal remote sensing images proposed in this invention are as follows:
[0077] 1. A boundary constraint change detection network BCnet based on a visual fundamental model is proposed. Using the visual fundamental model SAM as the backbone feature extractor, it fully mines the deep semantic information in dual-temporal remote sensing images. Based on this, a difference detail enhancement module is designed to strengthen the discriminative features between changed and unchanged regions. A multi-scale edge enhancement module is designed to finely model the change information between multi-temporal features, significantly improving the model's ability to perceive subtle changes and complex boundary details. An edge feature aggregation module is designed, which adaptively injects boundary information into change features through direction-aware attention modeling and residual fusion mechanisms, strengthening the spatial consistency and structural constraints of changed regions.
[0078] 2. An edge feature constraint strategy was studied, which provides dual guidance for the change region in the feature fusion stage and the feature output stage, strengthens the boundary constraint from both the feature space and the prediction space, and effectively suppresses the boundary ambiguity problem in change detection. Attached Figure Description
[0079] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings:
[0080] Figure 1 This is a schematic diagram of the boundary-constrained remote sensing image change detection model based on a visual fundamental model according to the present invention.
[0081] Figure 2 This is a schematic diagram of the structure of the difference detail enhancement module of the present invention;
[0082] Figure 3 This is a schematic diagram of the structure of the multi-scale edge enhancement module of the present invention;
[0083] Figure 4 This is a schematic diagram of the edge feature aggregation module of the present invention;
[0084] Figure 5 This is a schematic diagram illustrating the principle of the edge feature constraint strategy of the present invention;
[0085] Figure 6 This is a schematic diagram illustrating the BCnet detection effect of the present invention. Detailed Implementation
[0086] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0087] This invention proposes an intelligent interpretation method for dual-temporal remote sensing images, comprising:
[0088] Step S1: Construct a boundary-constrained remote sensing image change detection model based on a visual fundamental model;
[0089] The Boundary-Constrained Remote Sensing Image Change Detection Network Based on Vision Foundation Models (BCnet) includes: a SAM encoder module, a difference detail enhancement module, a multi-scale edge enhancement module, an edge feature aggregation module, and a change detection head. The edge feature aggregation module includes the edge detection head; the change detection head consists of a 3×3 convolutional kernel.
[0090] like Figure 1 As shown, the dual-temporal remote sensing image is input into the SAM encoder module to extract shallow semantic features and four levels of semantic features; the four levels of semantic features are input into the difference detail enhancement module to obtain four difference features; the shallow semantic features and the four difference features are input into the multi-scale edge enhancement module to obtain multi-scale edge features; on one hand, the multi-scale edge features are input into the edge detection head to obtain the edge mask; on the other hand, the multi-scale edge features are input into the edge feature aggregation module to obtain the edge aggregation features, and the edge aggregation features are input into the change detection head to obtain the change mask.
[0091] 1. SAM encoder module
[0092] SAM (Segment Anything Model), a fundamental visual model designed for general image segmentation tasks, originally only supported single-image input, making it difficult to directly adapt to the dual-temporal feature modeling requirements of intelligent interpretation of dual-temporal remote sensing images. Therefore, in the SAM encoder module, the single-input SAM image encoder is improved into a twin structure with two dual-temporal remote sensing image inputs. That is, the dual-temporal remote sensing images T1 and T2 are input to the first and second SAM image encoders respectively, and the two SAM image encoders share weights. Furthermore, to improve the adaptability of the SAM image encoder to remote sensing images, a LoRA adapter is introduced for efficient parameter fine-tuning, reducing training costs and enhancing relevance to change detection tasks while maintaining the original semantic representation capabilities.
[0093] 2. Difference Detail Enhancement Module
[0094] To highlight the variation information between the two-phase images, BCnet designed a difference detail enhancement module after the SAM encoder module to effectively mine the semantic differences between the two-phase images and enhance the discriminative power of the variation regions.
[0095] The difference detail enhancement module comprises four difference detail enhancement units (DDEUs). These four units (i.e., the first, second, third, and fourth DDEUs) are respectively connected to the 6th, 12th, 18th, and 24th Transformer layers in the SAM image encoder, which have the global receptive field. Specifically, the remote sensing image T1 is input into the first SAM image encoder, from which the difference detail enhancement units (DDEUs) are connected to the 6th, 12th, 18th, and 24th Transformer layers in the SAM image encoder, respectively. Semantic features are obtained from each Transformer layer. The remote sensing image T2 is input into the second SAM image encoder, from its corresponding... Semantic features are obtained from a Transformer layer. .
[0096] Subsequently, feature pairs and The first input DDEUs and generate output. .in, Indicates the level number of the Transformer. This indicates the cell number of the DDEU. The specific correspondence is as follows: and As the input of the first DDEU, the output ;Will and As the input to the second DDEU, the output ;Will and As the input of the third DDEU, the output ;Will and As the input of the fourth DDEU, the output .
[0097] To avoid unnecessary detail, the internal structure of the DDEU will be consistently described using the term "DDEU". and Representing the two input features of DDEU, using It indicates its output characteristics, without specifying a specific cell number.
[0098] like Figure 2 As shown, DDEU contains two parallel branches: a detail guidance branch and a difference modeling branch. In the detail guidance branch, the input features are... and ( The splicing is performed along the channel dimension to obtain the spliced features. This operation can preserve the most detailed information from both time phases. That is:
[0099]
[0100] in, Indicates splicing.
[0101] Subsequently, the spliced features are processed through two convolutional operations. Enhancement and compression are performed: the first convolutional layer is used to enhance features, and the second convolutional layer is used to reduce the number of channels to a smaller dimension. , obtain detailed guided features .
[0102] In the differential modeling branch, the element-wise absolute difference between the two input features is first calculated to generate differential features. :
[0103]
[0104] To simultaneously capture the distribution characteristics of change information in both spatial and channel dimensions, The data is fed into two parallel sub-paths: on the channel sub-path, channel statistical features are obtained by global max pooling along the spatial axis. Spatial response features are obtained by compressing along the channel axis using global max pooling on the spatial sub-path.
[0105] Subsequently, channel attention weights and spatial attention weights are generated through convolution operations and a sigmoid activation function, respectively. and :
[0106]
[0107] in, This represents the Sigmoid function, used to constrain weights within the interval [0,1].
[0108] Finally, the aforementioned attention weights are applied to the detail-guided features respectively. Enhanced differential features are generated through weighted fusion. :
[0109]
[0110] Through the above design, DDEU enhances the changing regions while maintaining the original semantic consistency, which can improve the model's ability to model small changing targets and complex boundary details, and provide high-quality change-aware feature representations for subsequent multi-scale edge enhancement.
[0111] 3. Multi-scale edge enhancement module
[0112] In high-resolution remote sensing image change detection, the boundaries of changed regions often exhibit multi-scale and irregular characteristics. To effectively capture these features and improve the detection accuracy of changed boundaries, this invention designs a Multi-Scale Edge Enhancement Module (MEEM). This module constructs a multi-level feature representation with strong robustness to changed boundaries through cross-layer feature alignment, hierarchical fusion, and edge-sensitive modeling.
[0113] like Figure 3 As shown, the input to MEEM includes semantic fusion features from the first Transformer layer of the SAM image encoder. and four difference features from the output of the difference detail enhancement module. , , , In this process, the remote sensing image T1 is input into the first SAM image encoder, which extracts shallow semantic features from the first Transformer layer, which has the global receptive field. The remote sensing image T2 is input into the second SAM image encoder to obtain shallow semantic features from its corresponding first Transformer layer. ; and Element-wise addition .
[0114] The processing flow of this module is as follows:
[0115] (1) Extraction of prior features of shallow edges
[0116] Given the powerful local structure and contour information modeling capabilities of the early Transformer layers in the SAM image encoder, MEEM first focuses on shallow semantic fusion features. Multi-scale convolution processing is performed. A set of parallel 3×3 convolutions combined with upsampling are used to transform the features into shallow edge priors with different spatial resolutions but uniform channel dimensions. Specifically, The first sub-path is obtained by using four parallel sub-paths, each of which undergoes an 8x upsampling and a 3×3 convolution. The second sub-path is obtained by 4x upsampling and 3×3 convolution. The third sub-path is obtained by 2x upsampling and 3×3 convolution. The fourth sub-path is obtained by upsampling by 1x and convolution by 3×3. .
[0117] (2) Alignment and enhancement of differential features
[0118] Difference features from the output of the difference detail enhancement module First, transpose convolution is used to enhance its scale information representation capability. Then, 1×1 convolution is used to compress and align the channel dimensions to obtain a multi-scale differential feature representation with a uniform number of channels. :
[0119]
[0120] in, Indicates transposed convolution; This represents convolution.
[0121] (3) Feature fusion
[0122] shallow edge prior features Multiscale differential features By adding dimensions, we obtain the fused features. :
[0123]
[0124] This fusion operation effectively compensates for the lack of spatial geometric details in deep semantic features, enabling features to simultaneously take into account semantic discriminative power and boundary accuracy.
[0125] (4) Two-way interaction
[0126] A two-way interactive strategy combining top-down and bottom-up approaches is adopted: high-level semantic features are downsampled and passed to the low-level layer to suppress background noise; low-level detail features are upsampled to supplement high-frequency edge information.
[0127] Specifically , By downsampling and Adding them together gives , By downsampling and Adding them together gives , By downsampling and Adding them together gives . , Through upsampling and Adding them together gives , Through upsampling and Adding them together gives , Through upsampling and Adding them together gives .
[0128] (5) Edge enhancement and multi-scale feature integration
[0129] After each level of fusion in the bidirectional interaction, a multilayer perceptron is introduced to perform nonlinear mapping and channel recalibration of the features to further enhance the boundary response and obtain edge-enhanced features at each scale. .Right now Obtained through a multilayer perceptron .
[0130] Edge enhancement features output at each scale The resolution is uniformly adjusted to 128×128, and the data is stitched together along the channel dimension to form the final multi-scale edge feature representation:
[0131]
[0132] Multi-scale edge features simultaneously integrate shallow, fine-grained boundary information with deep semantic change cues, providing high-quality boundary-aware feature support for subsequent edge detection heads and change detection heads.
[0133] 4. Edge Feature Aggregation Module
[0134] To enable effective feedback and guidance of explicit boundary information on changing features, an Edge Feature Aggregation Module (EFAM) was designed. Through direction-aware attention modeling and residual fusion mechanism, boundary information is adaptively injected into changing features, thereby strengthening the spatial consistency and structural constraints of changing regions.
[0135] like Figure 4 As shown, EFAM uses the output features of MEEM As input, EFAM consists of an edge detection head and two convolutional layers; The path taken by the edge detection head is called the edge detection head path. The path that goes directly through two convolutional layers is called the original feature path.
[0136] To effectively model the directional distribution characteristics of the boundary, EFAM designed a dual-branch directional awareness attention structure in the edge detection head.
[0137] Specifically, in the edge detection head, the input features are... Max pooling is performed along both the horizontal (row) and vertical (column) directions to obtain statistical features sensitive to both directions. and This operation not only captures long-range contextual dependencies in the horizontal direction but also preserves positional relationships in the vertical direction, helping the model to more accurately locate the spatial extent of changing objects.
[0138] Then, and The data are fed into convolutional layers, and normalized spatial horizontal attention weights are generated using the Sigmoid activation function. Attention weights in the perpendicular direction of space .
[0139]
[0140] in, This represents the Sigmoid activation function; This represents convolution.
[0141] Attention weights reflect the probability of significant boundary responses in corresponding rows or columns, thus enabling precise localization of changing regions. After obtaining the attention weights, EFAM adaptively reweights the input features through element-wise multiplication to achieve refined focusing on changing regions. As shown below:
[0142]
[0143] To further enhance the consistency and discriminative power of local features, EFAM reweights the features. Apply a convolutional transformation.
[0144] Meanwhile, in the original feature path, double convolution pairs are used. Nonlinear projection is performed to ensure semantic alignment between the original feature path and the edge detection head path. To prevent excessive boundary guidance from corrupting the original multi-scale edge features, the module employs a residual connection mechanism to fuse the edge detection head path with the original feature path to obtain aggregated edge features. As shown below:
[0145]
[0146] This design enhances the clarity of the boundaries and spatial positioning capabilities of the changed areas, providing highly discriminative input features for subsequent change detection heads.
[0147] Step S2: Based on the remote sensing image dataset, train the change detection model to obtain a trained change detection model;
[0148] In this embodiment, the publicly available remote sensing image dataset LEVIR-CD is used and proportionally divided into training, validation, and test sets. This dataset focuses on monitoring building changes over a long period, containing 637 pairs of 0.5-meter resolution dual-temporal images spanning 16 years (2002–2018). The images cover various complex building examples in Texas, USA, such as high-rise apartments, villas, and garages. The main challenges lie in the differences in lighting conditions and the diversity of building textures.
[0149] The training set is input into the change detection model BCnet for training.
[0150] In intelligent interpretation of dual-temporal remote sensing images, accurate localization of changed regions relies not only on strong semantic discrimination capabilities but also on the clarity of boundary structures. However, traditional training methods often only supervise based on masks of changed regions, neglecting explicit geometric constraints on the changed boundaries. This deficiency can easily lead to problems such as blurred boundaries, jagged edges, and structural breaks in the prediction results. To address this, this invention proposes an Edge Feature Constraint Strategy (EFCS), which constructs high-quality boundary supervision signals and jointly optimizes them with the change detection task, imposing stable and effective boundary constraints on the model during the training phase.
[0151] like Figure 5 As shown, EFCS consists of three parts: edge label generation, edge supervision loss, and edge back-guiding.
[0152] (1) Edge label generation
[0153] Given that existing remote sensing image datasets often lack explicit boundary annotations, this invention designs a morphology-based automatic edge label generation strategy. A reliable boundary supervision signal is constructed from the ground truth labels (binary) using an edge generator. Specifically, firstly, the ground truth labels of changing regions are smoothed using Gaussian filtering to suppress noise and alleviate the discreteness of the labeled edges. Then, the initial boundaries of the changing regions are extracted from the smoothed ground truth labels using the Canny edge detection operator. To enhance the continuity and fault tolerance of the boundaries, a dilation operation is performed on the initial boundaries to obtain the edge labels. This generation strategy can obtain edge labels with clear structure and a certain width without introducing additional manual annotation costs.
[0154] (2) Edge supervision loss
[0155] During training, BCnet incorporates multi-scale edge features. Input edge detection head branch, output boundary prediction probability map (i.e., edge mask). To balance pixel-level classification accuracy with overall shape structural consistency, an edge loss function combining binary cross-entropy loss (BCE Loss) and Dice loss is adopted. :
[0156]
[0157] in, This represents the binary cross-entropy loss; This indicates Dice's loss. Used to constrain the point-by-point prediction accuracy of boundary pixels, while It can effectively alleviate the extreme class imbalance between edge pixels and background pixels, and enhance the coherence and integrity of the prediction boundary in the overall structure.
[0158] (3) Edge reverse guidance
[0159] In order to truly involve boundary information in the feature learning process of change detection, EFCS is calculated not only at the output of the edge detection head branch. Furthermore, the learned edge aggregation features are also processed through the EFAM module. Reverse injection of change detection head branch, i.e., edge aggregation features The change prediction result (change mask) is input together with the boundary prediction probability map into the change detection head branch to obtain the change prediction result, thus achieving feature-level guidance. The change prediction result output by the change detection head branch... A binary cross-entropy loss is used to monitor the changing regions, as shown below:
[0160]
[0161] in, This represents the true result of the change (i.e., the true label).
[0162] Through the reverse guidance mechanism of edge aggregation features, the model can effectively improve the boundary clarity and spatial consistency of the predicted change detection results while maintaining semantic discriminative ability.
[0163] Combining the change detection task and the boundary constraint task, this invention uses the following joint loss function to train the model:
[0164]
[0165] in, This represents a trade-off coefficient used to balance the accuracy of change detection with the strength of boundary constraints. In this embodiment, Set it to 0.3 to ensure that the change detection and boundary constraint tasks are optimized together during training.
[0166] By employing an edge feature constraint strategy, BCnet achieves tight coupling between semantic discrimination of changed regions and boundary structure modeling during the training phase. This effectively alleviates the common problems of boundary blurring and lack of detail in change detection in complex scenarios, providing a stable and effective training mechanism for high-precision remote sensing image change detection.
[0167] Step S3: Input remote sensing images from different time phases into the trained change detection model and obtain the interpretation results.
[0168] The test set is input into the trained change detection model BCnet, and Precision, Recall, Intersection over Union (IoU), Overall Accuracy (OA), and F1 score are selected as evaluation metrics.
[0169] Among them, Precision focuses on the accuracy of the model's predictions of positive examples, Recall emphasizes the model's ability to capture positive examples, IoU is a comprehensive metric that measures the overlap between the predicted and actual regions of change, OA evaluates the overall accuracy of the model's classification, and F1 score balances the relationship between Precision and Recall. These evaluation metrics are calculated as follows:
[0170]
[0171]
[0172]
[0173]
[0174]
[0175] in, This represents the number of pixels that were predicted to change and actually did change; This represents the number of pixels that are predicted to remain unchanged and actually remain unchanged. This represents the number of pixels that were predicted to change but actually remained unchanged. This represents the number of pixels that were predicted to remain unchanged but actually changed. Figure 6 This is a partial schematic diagram of the BCnet detection results.
[0176] The average test results obtained are shown below:
[0177]
[0178] Experimental results confirm that BCnet has high robustness and accuracy advantages in challenging scenarios such as dense building clusters and complex edge cases.
[0179] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for intelligent interpretation of dual-temporal remote sensing images, characterized in that, include: Step S1: Construct a boundary-constrained remote sensing image change detection model based on a visual fundamental model; Step S2: Based on the remote sensing image dataset, train the change detection model to obtain a trained change detection model; Step S3: Input remote sensing images from different time phases into the trained change detection model and obtain the interpretation results; The boundary-constrained remote sensing image change detection model BCnet based on the visual basic model includes: a SAM encoder module, a difference detail enhancement module, a multi-scale edge enhancement module, an edge feature aggregation module, and a change detection head; wherein, the edge feature aggregation module includes an edge detection head; the change detection head includes a convolution with a 3×3 kernel. The dual-temporal remote sensing images are input into the SAM encoder module to extract shallow semantic features and four levels of semantic features; the four levels of semantic features are input into the difference detail enhancement module to obtain four difference features; the shallow semantic features and the four difference features are input into the multi-scale edge enhancement module to obtain multi-scale edge features; on one hand, the multi-scale edge features are input into the edge detection head to obtain the edge mask; on the other hand, the multi-scale edge features are input into the edge feature aggregation module to obtain the edge aggregation features, and the edge aggregation features are input into the change detection head to obtain the change mask. In the SAM encoder module, the single-input SAM image encoder is improved into a twin structure with two dual-temporal remote sensing image inputs. That is, the dual-temporal remote sensing images T1 and T2 are respectively input into the first SAM image encoder and the second SAM image encoder, and the two SAM image encoders share weights. The difference detail enhancement module comprises four difference detail enhancement units (DDEUs). These four units, namely the first, second, third, and fourth DDEUs, are connected to the 6th, 12th, 18th, and 24th Transformer layers in the SAM image encoder, respectively, which have the global receptive field. Specifically, the remote sensing image T1 is input into the first SAM image encoder, and the difference detail enhancement units (DDEUs) are connected to the 6th, 12th, 18th, and 24th Transformer layers in its global receptive field. Semantic features are obtained from each Transformer layer. The remote sensing image T2 is input into the second SAM image encoder, from its corresponding... Semantic features are obtained from a Transformer layer. ;in, ; Subsequently, feature pairs and The first input DDEUs and generate output. ;in, Indicates the level number of the Transformer. Indicates the cell number of the DDEU, and The specific correspondence is as follows: [The rest of the text appears to be a list of characters and symbols, possibly related to a specific project and As the input of the first DDEU, the output ;Will and As the input to the second DDEU, the output ;Will and As the input of the third DDEU, the output ;Will and As the input of the fourth DDEU, the output ; When describing the structure of DDEU, use and Representing the two input features of DDEU, using To represent its output features, DDEU contains two parallel branches: a detail guidance branch and a difference modeling branch; In the detailed guidance branch, the input features... and The splicing is performed along the channel dimension to obtain the spliced features. ,in, ;Right now: in, Indicates splicing; Subsequently, the spliced features are processed through two convolutional operations. Enhancement and compression are performed: the first convolutional layer is used to enhance features, and the second convolutional layer is used to reduce the number of channels to a smaller dimension. , obtain detailed guided features ; In the differential modeling branch, the element-wise absolute difference between the two input features is first calculated to generate differential features. : To simultaneously capture the distribution characteristics of change information in both spatial and channel dimensions, The data is fed into two parallel sub-paths: on the channel sub-path, channel statistical features are obtained by global max pooling along the spatial axis. Spatial response features are obtained by compressing along the channel axis using global max pooling on the spatial sub-path. ; Subsequently, channel attention weights and spatial attention weights are generated through convolution operations and a sigmoid activation function, respectively. and : in, This represents the Sigmoid function, used to constrain weights within the [0,1] interval; Finally, the aforementioned attention weights are applied to the detail-guided features respectively. Enhanced differential features are generated through weighted fusion. : 。 2. The method according to claim 1, characterized in that, The input to the Multi-Scale Edge Enhancement Module (MEEM) includes semantic fusion features from the first Transformer layer of the SAM image encoder. and four difference features from the output of the difference detail enhancement module. , , , In this process, the remote sensing image T1 is input into the first SAM image encoder, which extracts shallow semantic features from the first Transformer layer, which has a global receptive field. The remote sensing image T2 is input into the second SAM image encoder to obtain shallow semantic features from its corresponding first Transformer layer. ; and Obtained by adding elements one by one ; The MEEM processing flow is as follows: (1) Extraction of prior features of shallow edges: MEEM first analyzes shallow semantic fusion features. Multi-scale convolution processing is performed: a set of parallel 3×3 convolutions combined with upsampling are used to convert them into shallow edge prior features with different spatial resolutions but uniform channel dimensions. Specifically, The first sub-path is obtained by using four parallel sub-paths, each of which undergoes an 8x upsampling and a 3×3 convolution. The second sub-path is obtained by 4x upsampling and 3×3 convolution. The third sub-path is obtained by 2x upsampling and 3×3 convolution. The fourth sub-path is obtained by upsampling by 1x and convolution by 3×3. ; (2) Alignment and enhancement of differential features: Difference features from the output of the difference detail enhancement module First, transpose convolution is used to enhance its scale information representation capability. Then, 1×1 convolution is used to compress and align the channel dimensions to obtain a multi-scale differential feature representation with a uniform number of channels. : in, Indicates transposed convolution; Represents convolution; (3) Feature fusion shallow edge prior features Multiscale differential features By adding dimensions, we obtain the fused features. : (4) Two-way interaction A bidirectional interactive strategy combining top-down and bottom-up approaches is adopted: high-level semantic features are downsampled and passed to the lower level to suppress background noise; low-level detail features are upsampled to supplement high-frequency edge information. Specifically , By downsampling and Adding them together gives , By downsampling and Adding them together gives , By downsampling and Adding them together gives ; , Through upsampling and Adding them together gives , Through upsampling and Adding them together gives , Through upsampling and Adding them together gives ; (5) Edge enhancement and multi-scale feature integration After each level of fusion in the bidirectional interaction, a multilayer perceptron is introduced to perform nonlinear mapping and channel recalibration of the features to further enhance the boundary response and obtain edge-enhanced features at each scale. ,Right now Obtained through a multilayer perceptron ; Edge enhancement features output at each scale The resolution is uniformly adjusted to 128×128, and the data is stitched together along the channel dimension to form the final multi-scale edge feature representation: 。 3. The method according to claim 2, characterized in that, The edge feature aggregation module EFAM uses the output features of MEEM. As input, EFAM includes an edge detection head and two convolutional layers; The path taken by the edge detection head is called the edge detection head path. The path that directly passes through two convolutional layers is called the original feature path; In the edge detection head, the input features Max pooling is performed along both the horizontal and vertical directions to obtain statistical features sensitive to both directions. and This operation not only captures long-range contextual dependencies in the horizontal direction, but also preserves positional relationships in the vertical direction, which helps the model to more accurately locate the spatial extent of changing objects. Then, and The data are fed into convolutional layers, and normalized spatial horizontal attention weights are generated using the Sigmoid activation function. Attention weights in the perpendicular direction of space ; in, This represents the Sigmoid activation function; Represents convolution; After obtaining the attention weights, EFAM adaptively reweights the input features through element-wise multiplication to achieve refined focusing on regions of change; as shown below: To further enhance the consistency and discriminative power of local features, EFAM reweights the features... Apply a convolutional transformation; Meanwhile, in the original feature path, double convolution pairs are used. Perform nonlinear projection to ensure that the original feature path is aligned with the edge detection head path in semantic space; To prevent excessive boundary guidance from damaging the original multi-scale edge features, the module employs a residual connection mechanism to fuse the edge detection head path with the original feature path to obtain aggregated edge features. As shown below: 。 4. The method according to claim 3, characterized in that, An edge feature constraint strategy, EFCS, was also designed, which includes three parts: edge label generation, edge supervision loss, and edge back-guidance. (1) Edge label generation A morphology-based automatic edge label generation strategy is designed, which constructs reliable boundary supervision signals from binary ground truth labels using an edge generator. Specifically, firstly, Gaussian filtering is applied to the ground truth labels of changing regions to smooth them, suppressing noise and mitigating the discreteness of labeled edges. Subsequently, the initial boundaries of changing regions are extracted from the smoothed ground truth labels using the Canny edge detection operator. To enhance the continuity and fault tolerance of the boundaries, an expansion operation is performed on the initial boundaries to obtain the edge labels. ; (2) Edge supervision loss During training, BCnet incorporates multi-scale edge features. Input edge detection head branch, output boundary prediction probability map This refers to the edge mask; to balance pixel-level classification accuracy with the overall shape's structural consistency, an edge loss function combining binary cross-entropy loss (BCE Loss) and Dice loss is used. : in, This represents the binary cross-entropy loss; Indicates Dice loss; (3) Edge reverse guidance Edge aggregation features The change prediction result, i.e., the change mask, is jointly input into the change detection head branch along with the boundary prediction probability map, thus providing guidance at the feature level. The change prediction result output by the change detection head branch... Binary cross-entropy loss is used to supervise the change region; as shown below: in, This represents the true result of the change, i.e., the true label; Combining the change detection task and the boundary constraint task, the model is trained using the following joint loss function: in, This represents a tradeoff coefficient used to balance the accuracy of change detection with the strength of boundary constraints.
5. The method according to claim 4, characterized in that, It is 0.
3.
6. The method according to any one of claims 1-5, characterized in that, A LoRA adapter is introduced to fine-tune the parameters of the SAM image encoder.
7. The method according to claim 6, wherein the remote sensing image dataset in step S2 is LEVIR-CD.