Ground feature change detection method based on multi-scale feature fusion and attention mechanism

By employing a multi-scale feature fusion and attention mechanism, the problems of radiometric differences and insufficient feature extraction in dual-temporal RGB image change detection are solved, achieving higher accuracy and robustness in ground feature change detection.

CN121746927AInactive Publication Date: 2026-03-27JINHUA VOCATIONAL TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning methods are susceptible to false alarms and missed detections in dual-temporal RGB image change detection due to radiometric differences, insufficient feature extraction, and inadequate fusion of spatial channel information.

Method used

Employing a multi-scale feature fusion and attention mechanism, this approach generates dual-temporal attention-enhanced features through radiation consistency correction, joint data augmentation, weight-sharing twin backbone network, multi-scale feature complementarity module, and spatial-channel dual-dimensional verification attention module, ultimately outputting the area of ​​ground feature change.

Benefits of technology

It improves the accuracy and robustness of change detection, overcomes environmental noise interference, and achieves more complete and accurate identification of changes in land features at multiple scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746927A_ABST
    Figure CN121746927A_ABST
Patent Text Reader

Abstract

The invention discloses a ground feature change detection method based on multi-scale feature fusion and an attention mechanism, and the method comprises the steps: carrying out the radiation consistency correction and joint data enhancement of a dual-time-phase RGB image, and obtaining a preprocessed dual-time-phase image; adopting a weight sharing twin backbone network to extract multi-scale features of the preprocessed dual-time-phase image; ablating redundant information of the multi-scale features through a multi-scale feature complementation module to obtain differentiated features; weighting the differentiated features by using a space-channel two-dimensional verification attention module to generate two-time-phase attention enhancement features; and carrying out difference calculation and classification on the dual-temporal attention enhancement features, and outputting a ground feature change region. According to the embodiment of the invention, the method can improve the precision and robustness of change detection, overcomes the interference of environment noise, and achieves the more complete and more precise recognition of the change of a multi-scale ground feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image analysis technology, specifically a method for detecting ground feature changes based on multi-scale feature fusion and attention mechanism. Background Technology

[0002] Ground feature change detection is one of the core technologies in the field of remote sensing image analysis, with significant application value in environmental monitoring, urban expansion assessment, and disaster dynamic analysis. Utilizing dual-temporal RGB images for change detection has become a mainstream method due to its convenient data acquisition and strong intuitiveness. However, existing deep learning methods still face many challenges when dealing with such data: First, radiometric differences (such as illumination and seasonal variations) between dual-temporal images can easily be misjudged as real ground feature changes, generating false alarms; second, traditional methods lack the ability to represent ground features at different scales (such as from large buildings to small roads) during the feature extraction stage, making it difficult to simultaneously ensure the integrity of changes and boundary accuracy; furthermore, most models fail to effectively integrate spatial and channel-dimensional contextual information, resulting in limited ability to distinguish subtle changes and complex backgrounds, leading to noise or missed detections in the results. Summary of the Invention

[0003] The purpose of this invention is to provide a ground feature change detection method based on multi-scale feature fusion and attention mechanism to overcome the shortcomings of the prior art, improve the accuracy and robustness of change detection, overcome environmental noise interference, and achieve more complete and accurate identification of multi-scale ground feature changes.

[0004] One embodiment of this application provides a method for detecting ground cover changes based on multi-scale feature fusion and attention mechanism, the method comprising: Radiometric consistency correction and joint data augmentation are performed on the dual-temporal RGB images to obtain preprocessed dual-temporal images; A weight-sharing twin backbone network is used to extract multi-scale features from the preprocessed dual-temporal images; The redundant information of the multi-scale features is eliminated by the multi-scale feature complementation module to obtain differentiated features; The differential features are weighted using a spatial-channel dual-dimensional verification attention module to generate dual-temporal attention-enhanced features; The differences in the dual-temporal attention enhancement features are calculated and classified to output the areas of land cover change.

[0005] Optionally, the step of performing radiometric consistency correction and joint data augmentation on the dual-temporal RGB images to obtain preprocessed dual-temporal images includes: Load the original RGB images at time T1 and time T2 respectively, convert the image data into floating-point tensor format, and generate the original dual-temporal image tensor; Radiometric consistency correction is performed on the original dual-temporal image tensor. The mean and variance of each channel of the T1 and T2 images are calculated using a linear regression algorithm. The channel values ​​of the T2 image are adjusted to align with the statistical features of the T1 image to generate the radiometrically corrected image tensor. Joint data augmentation is performed on the radiometrically corrected image tensor by sequentially performing random rotation, horizontal flipping and scaling operations, while color dithering is applied to simulate illumination changes to generate the enhanced image tensor. The enhanced image tensor is uniformly scaled to a specific pixel size and then subjected to channel-level normalization to finally generate a preprocessed dual-temporal image.

[0006] Optionally, the step of using a weight-sharing Siamese backbone network to extract multi-scale features from the preprocessed dual-temporal images includes: A weight-sharing twin backbone network is constructed, and a four-stage feature extraction process is designed based on a lightweight ResNeSt50 architecture to generate a twin network model. The preprocessed dual-temporal images are input into the two branches of the Siamese network, and pixel-level edge features and color distribution features are extracted through the first-stage convolutional layer to generate a shallow feature map. Local texture features are extracted by combining a second-stage convolutional layer with max pooling, while grouped convolution is used to enhance feature diversity and generate a mid-level texture feature map. By introducing dilated convolutions with dilation rates of 2, 4, and 6 in the third-stage convolutional layer, the receptive field is expanded, global semantic features are extracted, and finally a feature pyramid containing multi-scale information is generated.

[0007] Optionally, the step of absolving redundant information of the multi-scale features through a multi-scale feature complementation module to obtain differentiated features includes: Feature maps from adjacent levels are selected from the multi-scale feature pyramid. The feature maps include shallow-middle and middle-deep feature combinations to generate a set of feature map pairs. For each feature map pair, perform element-wise subtraction to calculate the difference matrix between features at adjacent scales, eliminate recurring texture patterns, and generate initial difference features. Local response normalization is performed on the initial difference features to suppress background noise and enhance the response intensity of areas with significant changes, thereby generating enhanced difference features. By fusing all enhanced difference features through a 1×1 convolutional layer, cross-scale differential information is preserved, and finally an optimized differential feature tensor is generated.

[0008] Optionally, the step of using a spatial-channel dual-dimensional verification attention module to weight the differentiated features and generate dual-temporal attention-enhanced features includes: Spatial attention mechanism is applied to the differential feature tensor, and spatial location importance weights are calculated using 7×7 convolution kernels to generate a spatial attention weight map; A channel attention mechanism is applied to the differential feature tensor. The importance score of each channel is calculated through global average pooling and fully connected layers to generate a channel attention weight vector. The spatial attention weight map and the channel attention weight vector are broadcast and added together, and a preliminary fused attention map is generated by the Sigmoid function; By combining the feature alignment verification mechanism, the cosine similarity of the attention features at time T1 and T2 is calculated. Weight penalties are applied to similarity regions that exceed the preset similarity threshold, a refined attention map is generated and weighted to the differentiated features, and finally the dual-temporal attention enhancement features are output.

[0009] Optionally, the step of performing difference calculation and classification on the dual-temporal attention enhancement features and outputting the land cover change region includes: Perform pixel-wise absolute value difference operation on the dual-temporal attention enhancement features to obtain the initial feature difference map; Spatial context features are extracted from the initial feature difference map using a 3×3 convolutional layer to enhance the continuous representation of the changing regions and generate a context-enhanced difference map. The context-enhanced difference map is input into the dual-branch classifier, and the pixel-level change probability is calculated through a fully connected layer to generate a preliminary change probability map; Adaptive threshold segmentation is applied to the preliminary change probability map, and morphological post-processing is used to eliminate isolated noise points, ultimately outputting an accurate binary map of land cover change areas.

[0010] Another embodiment of this application provides a land cover change detection system based on multi-scale feature fusion and attention mechanism, the system comprising: The processing module is used to perform radiometric consistency correction and joint data augmentation on the dual-temporal RGB images to obtain preprocessed dual-temporal images; The extraction module is used to extract multi-scale features of the preprocessed dual-temporal image using a weight-shared twin backbone network. The ablation module is used to ablate redundant information of the multi-scale features through the multi-scale feature complementarity module to obtain differentiated features; The weighting module is used to weight the differential features using the spatial-channel dual-dimensional verification attention module to generate dual-temporal attention-enhanced features; The output module is used to perform difference calculation and classification on the dual-temporal attention enhancement features and output the land cover change area.

[0011] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0012] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0013] Compared with existing technologies, this invention provides a method for detecting ground cover changes based on multi-scale feature fusion and attention mechanisms. It performs radiometric consistency correction and joint data enhancement on dual-temporal RGB images to obtain preprocessed dual-temporal images. A weight-sharing Siamese backbone network is used to extract multi-scale features from the preprocessed dual-temporal images. A multi-scale feature complementarity module ablates redundant information of the multi-scale features to obtain differentiated features. A spatial-channel dual-dimensional verification attention module is used to weight the differentiated features, generating dual-temporal attention-enhanced features. Difference calculation and classification are performed on the dual-temporal attention-enhanced features to output the ground cover change areas. This improves the accuracy and robustness of change detection, overcomes environmental noise interference, and achieves more complete and accurate identification of multi-scale ground cover changes. Attached Figure Description

[0014] Figure 1 A hardware structure block diagram of a computer terminal for a method of detecting ground cover changes based on multi-scale feature fusion and attention mechanism provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for detecting ground cover changes based on multi-scale feature fusion and attention mechanism, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a ground feature change detection system based on multi-scale feature fusion and attention mechanism, provided in an embodiment of the present invention. Detailed Implementation

[0015] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0016] This invention first provides a method for detecting ground feature changes based on multi-scale feature fusion and attention mechanism. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0017] The following detailed explanation uses a computer terminal as an example. Figure 1This is a hardware structure block diagram of a computer terminal for a method of detecting ground cover changes based on multi-scale feature fusion and attention mechanism, provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0018] The non-volatile storage medium can store an operating system and a computer program. This computer program includes program instructions that, when executed, enable the processor to perform any ground feature change detection method based on multi-scale feature fusion and attention mechanisms.

[0019] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0020] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any ground feature change detection method based on multi-scale feature fusion and attention mechanism.

[0021] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0022] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0023] See Figure 2 The embodiments of the present invention provide a method for detecting ground cover changes based on multi-scale feature fusion and attention mechanism, which may include the following steps: S201, Perform radiometric consistency correction and joint data augmentation on the dual-temporal RGB image to obtain the preprocessed dual-temporal image; Specifically, the original RGB images at time T1 and time T2 can be loaded separately, the image data can be converted into floating-point tensor format, and the original dual-temporal image tensor can be generated. This step is the basic data preparation stage of preprocessing. The core is to ensure that the dual-temporal image data is in a uniform and computable format, providing standardized input for subsequent radiometric correction and enhancement. The specific implementation method is as follows: The original RGB images at time T1 and T2 are derived from drone aerial photography data of the same observation area. For example, T1 is an image of a suburban area of ​​the city collected in January 2024, and T2 is a re-measured image of the same area in July 2024. Both images are in lossless PNG format with a resolution of 1024×1024 pixels, containing three channels: red (R), green (G), and blue (B). The pixel value of each channel ranges from 0 to 255 (8-bit unsigned integer), corresponding to the light reflection intensity in the real scene (0 represents no reflection, and 255 represents maximum reflection). The loading process is completed through an image reading interface, automatically parsing metadata such as the image width, height, and number of channels to ensure that the size and channel structure of the two-phase images are completely consistent, avoiding subsequent processing failures due to data mismatch.

[0024] The core of data format conversion is to convert 8-bit unsigned integers (uint8) into 32-bit floating-point tensors (float32) while performing normalization to map pixel values ​​to the 0-1 range. This meets the input requirements of deep learning models and reduces the risk of numerical overflow. The conversion formula is: float_pixel = uint8_pixel / 255.0. For example, in image T1, the pixel values ​​of a building area are R=255, G=240, B=220, which are converted to (1.0, 0.941, 0.863); in image T2, the same location has increased pixel values ​​due to vegetation cover, resulting in R=30, G=180, B=50, which are converted to (0.118, 0.706, 0.196).

[0025] The generated original bi-temporal image tensor adopts a four-dimensional structure, with the dimension order being [number of phases, height, width, number of channels], i.e., [2, 1024, 1024, 3]. The first dimension "2" corresponds to the two time points T1 and T2, while the latter three dimensions correspond to the spatial size and channel information of a single image. The tensor is stored in the order of "T1 first, T2 last," with each element being a 32-bit floating-point number, retaining six decimal places to ensure the accuracy of subsequent numerical calculations. For example, the tensor fragment shows that the pixel values ​​(512, 512) at time T1 (index 0) are (1.0, 0.941, 0.863), and the pixel values ​​(512, 512) at time T2 (index 1) are (0.118, 0.706, 0.196), clearly presenting the pixel differences between the two temporal phases.

[0026] Radiometric consistency correction is performed on the original dual-temporal image tensor. The mean and variance of each channel of the T1 and T2 images are calculated using a linear regression algorithm. The channel values ​​of the T2 image are adjusted to align with the statistical features of the T1 image to generate the radiometrically corrected image tensor. The core objective of this step is to eliminate radiation distortion caused by differences in illumination intensity and imaging angle in dual-temporal images, and to avoid misjudging spurious changes (such as the difference in brightness between sunny and cloudy days) as real changes in ground features. The specific implementation method is as follows: The key to radiometric consistency correction is aligning the channel statistical features of images T1 and T2 using a linear regression algorithm. The statistical features selected are the mean (reflecting overall brightness) and variance (reflecting the dispersion of brightness distribution). First, the original dual-temporal image tensors are traversed, and the mean μ and variance σ of each channel in T1 and T2 are calculated respectively. The calculation range covers all pixels in the image, and the formula is μ_c = (1 / (H×W))×ΣΣpixel_c(i,j), σ_c 2 =(1 / (H×W))×ΣΣ(pixel_c(i,j)-μ_c) 2 Where c represents the R, G, and B channels, and H=1024 and W=1024 are the image dimensions.

[0027] For example, the statistical characteristics of image T1 are calculated as follows: R channel μ1_R=0.485, σ1_R=0.229; G channel μ1_G=0.456, σ1_G=0.224; B channel μ1_B=0.406, σ1_B=0.225. Image T2, due to being captured on a cloudy day with low brightness, has the following statistical characteristics: R channel μ2_R=0.390, σ2_R=0.201; G channel μ2_G=0.365, σ2_G=0.198; B channel μ2_B=0.320, σ2_B=0.190.

[0028] The core of the linear regression algorithm is to construct a mapping relationship from the T2 channel value to the T1 channel value. The mapping model is T2'_c = a_c × T2_c + b_c, where T2'_c is the corrected T2 channel value, and a_c and b_c are the regression coefficients, which are calculated through variance alignment and mean alignment respectively: a_c = σ1_c / σ2_c (to ensure that the brightness distribution of the two images is consistent), b_c = μ1_c - a_c × μ2_c (to ensure that the overall brightness of the two images is consistent). Substituting the above statistical characteristics, the regression coefficients are calculated as follows: R channel a_R=0.229 / 0.201≈1.139, b_R=0.485-1.139×0.390≈0.485-0.444≈0.041; G channel a_G=0.224 / 0.198≈1.131, b_G=0.456-1.131×0.365≈0.456-0.413≈0.043; B channel a_B=0.225 / 0.190≈1.184, b_B=0.406-1.184×0.320≈0.406-0.379≈0.027.

[0029] The pixel values ​​of each channel of the T2 image are substituted into the mapping model for correction. For example, for a pixel in the T2 image with R=0.118, after correction, T2'_R=1.139×0.118+0.041≈0.134+0.041≈0.175; G=0.706, after correction, T2'_G=1.131×0.706+0.043≈0.798+0.043≈0.841; B=0.196, after correction, T2'_B=1.184×0.196+0.027≈0.232+0.027≈0.259. After correction, the statistical characteristics of the T2 image are recalculated: μ2'_R≈0.484, σ2'_R≈0.228, and the deviation from the statistical characteristics of T1 is ≤0.002, achieving radiometric consistency alignment. The generated radiometrically corrected image tensor still maintains the four-dimensional structure of [2, 1024, 1024, 3]. The T1 image remains unchanged, while the T2 image contains the corrected data, effectively eliminating spurious changes caused by illumination differences.

[0030] Joint data augmentation is performed on the radiometrically corrected image tensor by sequentially performing random rotation, horizontal flipping and scaling operations, while color dithering is applied to simulate illumination changes to generate the enhanced image tensor. This step expands the distribution range of training samples through diverse data augmentation operations, improving the model's robustness to changes in shooting angle and lighting intensity. The core is to ensure that the augmentation operations on the two temporal images are executed synchronously to maintain spatial correspondence. The specific implementation method is as follows: All joint data augmentation operations are performed synchronously on the two temporal images, meaning that images T1 and T2 use the exact same augmentation parameters to avoid disrupting the spatiotemporal correspondence between the two images. First, a random rotation operation is performed, with the rotation angle randomly selected between -15° and 15° (a positive angle indicates clockwise rotation, and a negative angle indicates counter-clockwise rotation). For example, if a random rotation angle of 8° is generated, both images T1 and T2 will be rotated 8° clockwise. During the rotation process, a bilinear interpolation algorithm is used to fill the blank areas created by the image rotation. The interpolation kernel size is 3×3, ensuring that the edges of the rotated image are smooth and free of obvious jagged distortion. For example, the edges of buildings remain continuous and intact after rotation, without breaks or blurring.

[0031] The random horizontal flip operation is performed with a probability of 0.5, meaning there is a 50% probability that both temporal images will be horizontally flipped (mirrored) simultaneously, and a 50% probability that they will not be flipped. During the flip, pixels are symmetrically swapped along the vertical midline of the image. For example, in image T1, the road pixels (512, 200) on the left are swapped with the pixels (512, 824) on the right. The same operation is performed synchronously in image T2 to ensure that the positions of ground features in both temporal images still correspond one-to-one after the flip. For example, before the flip, the building in T1 is on the left and the newly added vegetation in T2 is on the right side of the building. After the flip, the building is on the right side and the newly added vegetation is still on the right side of the building, and the spatial relationship remains unchanged.

[0032] The scaling factor for the scaling operation is randomly selected between 0.8 and 1.2. For example, a randomly generated scaling factor of 1.1 will enlarge both temporal images by a factor of 1.1, scaling them from 1024×1024 pixels to 1126×1126 pixels, and then restoring them to 1024×1024 pixels through center cropping. If the scaling factor is 0.9, the image will be reduced to 921×921 pixels, and then restored to 1024×1024 pixels by filling the edges with black pixels (pixel value 0). The scaling process also uses bilinear interpolation to ensure that texture features are not lost; for example, the leaf texture of vegetation and the details of road markings remain clearly distinguishable after scaling.

[0033] Color dithering is used to simulate lighting changes in real-world scenes while maintaining the relative differences between the two phases. Adjustment parameters include brightness (±10%), contrast (±15%), and saturation (±20%). Brightness adjustment is achieved by multiplying the pixel values ​​of all channels by an adjustment factor; for example, an 8% increase in brightness results in a uniform multiplication of pixel values ​​by 1.08. Contrast adjustment is achieved by stretching the distribution range of pixel values; for example, a 12% increase in contrast results in pixel value x being adjusted to x×1.12+(0.5-0.5×1.12). Saturation adjustment is achieved by adjusting the relative proportions of the RGB channels; for example, a 15% decrease in saturation results in a 15% increase in the weight of each pixel's grayscale value ((R+G+B) / 3) and a 5% decrease in the weight of each channel. For example, pixels (0.175, 0.841, 0.259) after T2 correction, after a 5% increase in brightness and a 10% increase in contrast, become (0.199, 0.961, 0.297). The same adjustments are performed synchronously on the T1 image to maintain the differences between the two phases. The generated enhanced image tensor is still [2, 1024, 1024, 3], containing diverse scene variations, providing rich samples for model training.

[0034] The enhanced image tensor is uniformly scaled to a specific pixel size and then subjected to channel-level normalization to finally generate a preprocessed dual-temporal image.

[0035] This step ensures that the image data input to the model has a consistent spatial scale and numerical distribution by unifying the size and standardizing, thereby improving the convergence speed and detection accuracy of model training. The specific implementation method is as follows: The target pixel size for uniform scaling is set to 512×512. This size balances model computational efficiency and feature preservation, and is suitable for the input requirements of the lightweight ResNeSt50 backbone network. The scaling process employs a bilinear interpolation algorithm, which obtains the target pixel value by weighted averaging of the values ​​of the four surrounding pixels. The formula is pixel_target = w1×pixel1 + w2×pixel2 + w3×pixel3 + w4×pixel4, where the weights w1-w4 are calculated based on the distance between the target pixel and the original pixel; the closer the distance, the greater the weight. For example, a 4×4 pixel block in the enhanced T2 image corresponds to a 1×1 pixel in the 512×512 image after scaling. This pixel value is the weighted average of the original four pixels, ensuring a smooth transition of texture features and avoiding loss of detail due to scaling. For instance, the straightness of road edges remains intact after scaling.

[0036] Channel-level normalization employs the Z-score normalization method. Based on pre-computed global statistical parameters from a large-scale RGB ground cover image dataset, the pixel values ​​of each channel are normalized to ensure they fall within the range of -1 to 1. The formula is pixel_norm = (pixel - μ_global) / σ_global, where μ_global is the global mean (R = 0.485, G = 0.456, B = 0.406) and σ_global is the global standard deviation (R = 0.229, G = 0.224, B = 0.225). These parameters are obtained by statistically analyzing hundreds of thousands of RGB ground cover images, making the distribution of the input data closer to the distribution during model training and improving generalization ability.

[0037] For example, the pixels (0.199, 0.961, 0.297) of the enhanced T2 image are normalized as follows: R channel pixel_norm_R = (0.199 - 0.485) / 0.229 ≈ (-0.286) / 0.229 ≈ -1.249; G channel pixel_norm_G = (0.961 - 0.456) / 0.224 ≈ 0.505 / 0.224 ≈ 2.254. Since the value exceeds the range of -1 to 1, it is truncated and limited to 1.0; B channel pixel_norm_B = (0.297 - 0.406) / 0.225 ≈ (-0.109) / 0.225 ≈ -0.484. The standardized pixel values ​​are (-1.249, 1.0, -0.484), which preserves the differences between channels and avoids the impact of extreme values ​​on the model.

[0038] The final preprocessed dual-temporal images are four-dimensional tensors with dimensions [2, 512, 512, 3]. The pixel values ​​of each channel in both T1 and T2 images conform to a standard normal distribution (mean ≈ 0, standard deviation ≈ 1). The elements in the tensor are preserved to four decimal places. For example, a pixel in the T1 image, after standardization, is (0.3215, 0.2890, 0.1987), while the corresponding pixel in the T2 image is (-1.2490, 1.0000, -0.4840). This clearly presents the numerical differences corresponding to changes in ground features, laying a solid foundation for subsequent multi-scale feature extraction.

[0039] S202, using a weight-sharing twin backbone network, extract multi-scale features of the preprocessed dual-temporal image; Specifically, a weight-sharing twin backbone network can be constructed, and a four-stage feature extraction process can be designed based on the lightweight ResNeSt50 architecture to generate a twin network model. This step is the core architecture design for multi-scale feature extraction. Through lightweight modifications and a weight-sharing mechanism, it reduces computational costs while ensuring feature extraction capabilities and guarantees consistent comparison of features from two temporal images. The specific implementation is as follows: The lightweight ResNeSt50 architecture is based on the original ResNeSt50 and optimizes the parameter scale for the characteristic of RGB images with only three channels: First, the number of output channels of the initial convolutional layer is reduced from the original 256 channels to 64 channels, reducing the computational cost of shallow layers; Second, a bottleneck structure is adopted in the residual block (1×1 convolution for dimensionality reduction → 3×3 grouped convolution → 1×1 convolution for dimensionality increase), and the number of groups of grouped convolution is adjusted from 8 groups to 4 groups, reducing the number of parameters while preserving feature diversity. Finally, the number of model parameters is controlled within 8M to meet the requirements of real-time processing.

[0040] The four feature extraction stages are designed to match the feature hierarchy of RGB images. Each stage contains several residual blocks and convolutional layers, with the specific configuration as follows: The first stage (shallow features) contains one convolutional layer (3×3 convolutional kernel, stride 2, padding=1, output channels 64) and two residual blocks. The residual blocks use 1×1 convolution (reduced to 16 channels), 3×3 grouped convolution (4 groups, 16 channels), and 1×1 convolution (increased to 64 channels). The activation function is ReLU for all of them. This stage focuses on low-level edge and color features. The second stage (intermediate texture) includes one max pooling layer (2×2 pooling kernel, stride 2), one convolutional layer (3×3 convolutional kernel, stride 1, padding=1, output channels 128) and three residual blocks. The residual block channel size is 128. The number of grouped convolutions is maintained at 4, focusing on extracting local texture structures. The third stage (global semantics) contains one convolutional layer (3×3 dilated convolution, dilation rate 2 / 4 / 6 in parallel, padding=2 / 4 / 6, output channels 256) and five residual blocks. The residual blocks are integrated into the dilated convolution to expand the receptive field and capture the semantic information of the "ground object-background" association. The fourth stage (deep fusion) contains one convolutional layer (3×3 convolutional kernel, stride 1, padding=1, output channels 512) and two residual blocks. It fuses the features from the previous stage through 1×1 convolution to output high-dimensional semantic features.

[0041] The weight-sharing mechanism is the core of Siamese networks. This means that the two branches (processing images T1 and T2 respectively) share the same set of convolutional kernel weights and residual block parameters. Only one set of weights is updated during training, ensuring consistent standards for feature extraction across both temporal phases and avoiding feature comparison biases caused by differences in branch parameters. For example, the first-stage convolutional kernel weight for the branch processing image T1 is W1, and the convolutional kernel weight at the same position in the branch processing image T2 is also W1. Both are updated synchronously during backpropagation. The resulting Siamese network model takes a pair of temporal images as input and outputs a multi-scale feature sequence from both branches.

[0042] The preprocessed dual-temporal images are input into the two branches of the Siamese network, and pixel-level edge features and color distribution features are extracted through the first-stage convolutional layer to generate a shallow feature map. This step captures the most intuitive change cues in RGB images through shallow convolution, focusing on edges and color—the two most sensitive features of RGB data—laying the foundation for subsequent texture and semantic extraction. The specific implementation is as follows: The preprocessed dual-temporal images are tensors of [2, 512, 512, 3], where index 0 corresponds to image T1 (e.g., a suburban image from January 2024, containing red buildings and gray roads), and index 1 corresponds to image T2 (the same area from July 2024, where some red buildings have been demolished and new green vegetation has been added). The two images are then input into branches 1 (T1 processing) and 2 (T2 processing) of the Siamese network, respectively. The convolutional layer parameters in the first stage of both branches are identical: 3×3 kernels, 64 kernels, stride 2, padding=1, ensuring that the input 512×512 image is reduced to 256×256 after convolution, and the number of channels increases from 3 to 64.

[0043] The extraction of pixel-level edge features depends on the response of the convolution kernel to changes in grayscale. For example, the vertical edge convolution kernel for the building edge (with a weight matrix of [[-1,0,1],[-1,0,1],[-1,0,1]]) generates a high response value (about 0.9) at the boundary between the red building area (approximately 0.8-1.0 after pixel value normalization) and the surrounding road area (0.3-0.5) in image T1, forming continuous edge features. In image T2, the original building area becomes vegetation (0.2-0.4), and the edge convolution kernel generates a high response (approximately 0.8) at the boundary between the vegetation and the road, but the edge shape changes from rectangular to irregular, clearly reflecting the edge changes after the building is demolished.

[0044] Color distribution characteristics are reflected through differences in convolutional responses of different channels. For example, the response value of the red channel (R) convolutional kernel to the T1 building area (0.9) is significantly higher than that to the T2 vegetation area (0.3), the response value of the green channel (G) convolutional kernel to the T2 vegetation area (0.8) is higher than that to the T1 building area (0.2), and the response value of the blue channel (B) to the road area in both time phases is stable between 0.4 and 0.5. These channel response differences are encoded into different dimensions of a 64-channel shallow feature map. For example, channel 10 (emphasizing the edge of the R channel) has a high value in the T1 building area, and channel 25 (emphasizing the edge of the G channel) has a high value in the T2 vegetation area.

[0045] The generated shallow feature maps are two branches, each outputting a 256×256×64 feature map. The T1 branch feature map is dominated by high responses to the edges of red buildings and gray roads, while the T2 branch feature map is dominated by high responses to the edges of green vegetation and the edges of remaining buildings. The difference regions (original building locations) between the two maps have shown preliminary feature comparisons, providing clear candidate regions for changes in subsequent mid-layer texture extraction.

[0046] Local texture features are extracted by combining a second-stage convolutional layer with max pooling, while grouped convolution is used to enhance feature diversity and generate a mid-level texture feature map. This step, based on shallow edge and color features, further mines local texture structures. By using grouped convolution to capture the unique textures of different types of land features in the RGB image, it improves the distinguishability of features. The specific implementation method is as follows: First, max pooling is performed with a 2×2 kernel size, a stride of 2, and padding=0. This reduces the size of the 256×256×64 feature map output from the first stage to 128×128, while preserving the maximum response values ​​in local areas. This enhances significant texture features (such as building brick patterns and vegetation veins) and suppresses weak noise (such as local color fluctuations caused by road stains). For example, after pooling, the brick pattern in the building area of ​​image T1 still clearly retains its "interlaced" texture, while the sporadic high responses of road stains are filtered out.

[0047] The second-stage convolutional layer uses 3×3 convolutional kernels, 128 in total, with a stride of 1 and padding of 1, ensuring the feature map size remains 128×128, and increasing the number of channels from 64 to 128. The core innovation lies in introducing grouped convolution, dividing the 128 convolutional kernels into 4 groups (32 kernels per group). Each group performs convolution operations only on 16 channels (64 channels ÷ 4 groups) of the input feature map, and then concatenates the outputs of the 4 groups to form a 128-channel feature map. The advantage of grouped convolution is that it allows different groups of convolutional kernels to focus on different texture types, for example: Group 1 convolutional kernels: capture horizontal textures (such as road markings and building waistlines), with a response value of 0.8 for the white horizontal markings of road T1; The second set of convolutional kernels captures vertical textures (such as building columns and tree trunks), achieving a vertical trunk response value of 0.9 for T2 vegetation. Group 3 convolutional kernels: capture mesh textures (such as architectural window panes and fences), with a mesh structure response value of 0.85 for T1 architectural window panes; Group 4 convolution kernels: capture irregular textures (such as vegetation clumps and land areas), with a response value of 0.75 for irregular textures of newly added vegetation in T2.

[0048] The extraction of local texture features directly reflects changes in land cover types. For example, the "brick pattern + window frame" texture combination in the building area of ​​T1 becomes the "leaf vein + trunk" vegetation texture combination in the T2 image. The response channels and intensities of the two types of textures in the feature map are significantly different: the 30-45 channels of the T1 feature map (output of the 2nd group of vertical texture kernels) have high values ​​(0.7-0.9), while the 60-75 channels of the T2 feature map (output of the 4th group of irregular texture kernels) have high values ​​(0.6-0.8).

[0049] The generated mid-level texture feature maps consist of two branches, each containing a 128×128×128 feature map. Compared to the shallow feature maps, these maps offer richer texture details and higher differentiation between land cover types. For example, the texture features of building areas and road areas are clearly distinguishable in the T1 branch feature map, and the texture differences between newly added vegetation areas and remaining building areas are also readily apparent in the T2 branch feature map, providing refined local feature support for subsequent global semantic extraction.

[0050] By introducing dilated convolutions with dilation rates of 2, 4, and 6 in the third-stage convolutional layer, the receptive field is expanded, global semantic features are extracted, and finally a feature pyramid containing multi-scale information is generated.

[0051] This step overcomes the receptive field limitations of conventional convolution through dilated convolution, capturing the global semantic associations of ground features in RGB images (such as the combined variations of "building-road" and "vegetation-plot"), while integrating multi-stage features to form a pyramid, covering full-scale information from edges to semantics. The specific implementation is as follows: The third-stage convolutional layer employs atrous convolution, the core of which is inserting "holes" (zero padding) between convolutional kernel elements to expand the receptive field without increasing parameters or computational cost. Three dilation rates (2, 4, 6) are designed for parallel operation of atrous convolutions, each corresponding to a set of 3×3 convolutional kernels (64 kernels each). The input is the 128×128×128 feature map output from the second stage, with the specific configuration as follows: Inflation rate 2: The effective receptive field of the convolution kernel is 5×5 (3+(3-1)×2), padding=2, ensuring the output size is 128×128, capturing "local relationships of ground features" (such as the boundary between buildings and adjacent roads). Inflation rate 4: Effective receptive field is 11×11(3+(3-1)×4), padding=4, capturing “moderate correlation of ground features” (such as a block composed of multiple buildings). Inflation rate 6: The effective receptive field is 15×15(3+(3-1)×6), padding=6, capturing “global correlation of ground features” (such as the distribution of blocks and surrounding vegetation).

[0052] The convolutional outputs of the three dilation rates (each 128×128×64) are concatenated and fused into a 128×128×192 feature map. This is then compressed through a 1×1 convolutional layer (output channels 256) to obtain a 128×128×256 global semantic feature map. The global semantic features focus on changes in the relationships between land cover combinations, for example: In the T1 image, the combination of "red building (128×128 area) + gray road (surrounding 200×200 area)" forms a high response (0.7-0.9) in the 100-150 channels of the semantic feature map, representing the "building-road" semantics; In image T2, the original building area is transformed into "green vegetation (128×128 area) + gray road (surrounding 200×200 area)". The semantic feature map channels 50-100 (vegetation-related semantics) show high responses (0.6-0.8), representing the "vegetation-road" semantics, which is significantly different from the semantic channels of T1.

[0053] The feature pyramid generation integrates the output features from the first four stages, and upsampling (bilinear interpolation) unifies the feature map size of each stage to 256×256, forming a four-scale feature set: First stage (shallow layer): 256×256×64 (edge ​​+ color); Second stage (middle layer): 256×256×128 (texture, upsampled by 2x); Third stage (deep layer): 256×256×256 (semantic, upsampled by 2 times); Fourth stage (fusion): 256×256×512 (1×1 convolution fusion of the first three stages, outputting 512 channels).

[0054] This feature pyramid covers multi-scale information from 64×64 pixels (shallow details) to 256×256 pixels (deep semantics). Each scale of features contains contrast information between T1 and T2 phases, such as the edge difference between T1 and T2 in the shallow pyramid and the semantic difference between T1 and T2 in the deep pyramid, providing full-dimensional feature input for subsequent feature complementation and attention weighting.

[0055] S203, the redundant information of the multi-scale features is eliminated by the multi-scale feature complementation module to obtain differentiated features; Specifically, feature maps from adjacent levels can be selected from the multi-scale feature pyramid. These feature maps include shallow-middle and middle-deep feature combinations to generate a set of feature map pairs. This step focuses on the differences and correlations of cross-scale features by filtering adjacent feature maps, laying the foundation for subsequent redundancy ablation. The core is to ensure the size consistency and information complementarity of feature map pairs. The specific implementation method is as follows: The multi-scale feature pyramid comprises four stages of features, with adjacent levels combined into "shallow-middle" and "middle-deep" layers. The spatial dimensions (the number of channels can differ) of the feature maps at each level must be standardized beforehand to avoid computational errors due to size differences. The original parameters for each level of the feature pyramid are as follows: Shallow feature map (first stage): size 256×256 pixels, number of channels 64, focusing on edge and color features (such as building outlines, road color boundaries). Mid-layer feature map (second stage): size 128×128 pixels, number of channels 128, focusing on local texture features (such as building brick texture, vegetation leaf veins). Deep feature map (third stage): 128×128 pixels in size, 256 channels, focusing on global semantic features (such as "building-road" and "vegetation-plot" associations).

[0056] Size uniformity is achieved through bilinear interpolation: the shallow feature map is upsampled from 256×256 to 128×128 (consistent with the middle and deep layers), with an interpolation kernel size of 3×3, ensuring smooth transitions in edge features. For example, the vertical edges of buildings in the shallow layer remain continuous after sampling, without breaks or blurring. After upsampling, the shallow feature map parameters become 128×128×64, perfectly matching the spatial dimensions of the middle layer (128×128×128) and the deep layer (128×128×256).

[0057] The construction of feature map sets follows the principle of "adjacent levels + complementary information": Shallow-Medium Feature Pair: Composed of upsampled shallow feature map (128×128×64) and medium feature map (128×128×128). The shallow layer provides low-dimensional edge / color reference, while the medium layer provides high-dimensional texture details. The two complement each other and can explore the difference between "edge and texture" (such as the building edge remains unchanged but the texture changes from brick texture to vegetation texture). The mid-layer-deep feature pair consists of a mid-layer feature map (128×128×128) and a deep feature map (128×128×256). The mid-layer provides local texture reference, while the deep layer provides global semantic association. The two complement each other and can explore the differences between "texture and semantics" (such as the semantic change of "bare land-vegetation" corresponding to the addition of local vegetation texture).

[0058] The generated feature map set contains two sets of feature pairs. The size of the two feature maps in each set is 128×128, and the number of channels is "64+128" and "128+256" respectively. Redundancy ablation will be performed on each set separately to ensure that cross-scale differences are not missed.

[0059] For each feature map pair, perform element-wise subtraction to calculate the difference matrix between features at adjacent scales, eliminate recurring texture patterns, and generate initial difference features. This step captures the differences in features between adjacent scales through element-wise subtraction. The core is to eliminate redundant information that repeats across scales (such as general textures and fixed edges) and highlight the feature differences caused by real changes. The specific implementation is as follows: The element-wise subtraction operation is performed for each pair of feature maps according to the principle of "channel matching + pixel alignment": For shallow-middle feature pairs, the 64 channels of the shallow feature map (128×128×64) correspond one-to-one with the first 64 channels (focusing on basic texture features) of the middle feature map (128×128×128), and the feature value at each pixel position (i,j) is subtracted by "middle feature value - shallow feature value"; For middle-deep feature pairs, the 128 channels of the middle feature map (128×128×128) correspond to the first 128 channels (focusing on basic semantic features) of the deep feature map (128×128×256), and the same pixel-wise subtraction is performed.

[0060] The physical meaning of the difference matrix is ​​"the magnitude of change in features at adjacent scales." A larger value indicates more significant differences in features across scales, potentially corresponding to actual changes in land cover. A smaller value indicates high repetition of features across scales, representing redundant information (e.g., the gray texture of roads shows similar responses in both shallow and mid-scale layers, resulting in a value close to 0 after subtraction). For example: In the shallow-to-medium layer feature pairing, the brick texture of the building area in image T1 (shallow channel 10, pixel (50,50) feature value 0.8; mid-layer channel 10, same pixel feature value 0.75) has a difference of 0.05 after subtraction (redundancy, no change); in image T2, the original building area becomes vegetation, shallow channel 10 pixel (50,50) feature value 0.3 (edge ​​disappearance), mid-layer channel 10 feature value 0.6 (new vegetation texture), and the difference of 0.3 after subtraction (significant difference, corresponding change). In the mid-to-deep feature pair, the semantic association of "building-road" in image T1 is 0.7 for the mid-level channel 60 pixel (80,80) and 0.72 for the deep channel 60 pixel, with a difference of 0.02 after subtraction (redundancy); in image T2, the association changes to "vegetation-road" at this location (0.5 for the mid-level channel 60 and 0.8 for the deep channel 60), with a difference of 0.3 after subtraction (significant difference).

[0061] The generated initial difference features are two sets of single-channel feature maps: shallow-to-medium layer difference features (128×128×64) and medium-to-deep layer difference features (128×128×128). All difference feature values ​​are restricted to non-negative values ​​by the ReLU activation function (only positive differences where "medium / deep features are stronger than shallow / medium features" are retained, and negative fluctuations caused by noise are excluded). For example, the difference value -0.03 becomes 0 after ReLU, while the difference value 0.3 remains unchanged, ensuring that the initial difference features only contain valid difference information.

[0062] Local response normalization is performed on the initial difference features to suppress background noise and enhance the response intensity of areas with significant changes, thereby generating enhanced difference features. This step optimizes the initial difference features through Local Response Normalization (LRN). The core idea is to suppress scattered background noise (such as random fluctuations in single pixels) while enhancing continuous areas of significant change (such as large areas of newly added vegetation). The specific implementation method is as follows: The core logic of local response normalization is to "normalize the feature value of the current pixel by using the mean and variance of the response of the local region," avoiding misjudging a high response of a single pixel as a change, while amplifying the intensity of continuous high response regions. The algorithm parameters are set as follows: normalization window size 5×5 (covering a 5×5 local region around the current pixel to balance local information and computational load), α=1e-4 (scaling factor to prevent excessively large values), β=0.75 (exponential factor to control the normalization intensity), k=2 (offset to avoid a denominator of 0), and the normalization formula is: LRN(x,i,j,c)=x(i,j,c) / [k+(α / n)×Σ(x(i',j',c))] 2]^β. Where x(i,j,c) is the feature value of the initial difference feature in channel c and pixel (i,j), n=25 (number of pixels in a 5×5 window), and Σ is the sum of squares of the feature values ​​of all pixels in the window.

[0063] The normalization process has a significantly different effect on noise and salient regions: Background noise region (such as the difference feature value of 0.1 caused by scattered stains on the road): other pixel values ​​within the window are mostly 0.01-0.05, the sum of squares ≈ 0.01, substituting into the formula, the LRN value is ≈ 0.1 / [2+(1e-4 / 25)×0.01]^0.75≈0.1 / 1.68≈0.06, the response intensity is reduced by 40%, and the noise is suppressed; Significantly changed areas (such as newly added vegetation areas in T2 images, with difference feature values ​​of 0.3-0.5): The pixel values ​​within the window are mostly 0.2-0.4, and the sum of squares is ≈2.5. Substituting into the formula, the LRN value is ≈0.4 / [2+(1e-4 / 25)×2.5]^0.75≈0.4 / 1.73≈0.23. Although the absolute value decreases, the contrast relative to the noise area is improved (from 0.1:0.01 to 0.23:0.06), and the significant areas are more prominent.

[0064] The generated enhanced difference features have the same size and number of channels as the initial difference features: shallow-to-medium layer enhanced difference features (128×128×64) and medium-to-deep layer enhanced difference features (128×128×128). The feature value distribution is more concentrated in the range of 0-0.5. The background noise ratio is reduced from the initial 25% to 10%, and the continuous response rate of the large-scale change area is increased from 60% to 85%, providing high-quality difference features for subsequent fusion.

[0065] By fusing all enhanced difference features through a 1×1 convolutional layer, cross-scale differential information is preserved, and finally an optimized differential feature tensor is generated.

[0066] This step achieves efficient fusion of multi-scale enhanced difference features through 1×1 convolution. The core is to preserve the complementarity of cross-scale differences while reducing dimensionality, avoiding information redundancy, and generating a unified differential feature representation. The specific implementation method is as follows: First, the channel dimension of the two sets of enhanced difference features is adapted: the shallow-to-medium layer enhanced difference feature (128×128×64) increases the number of channels from 64 to 128 through a 1×1 convolutional layer (128 convolutional kernels, stride 1, padding=0). The convolutional kernel weights learn the mapping relationship between "edge-texture difference" through training. For example, the shallow edge difference (channel 10) and the medium texture difference (channel 20) are merged into the feature of the new channel 5. The medium-to-deep layer enhanced difference feature (128×128×256) decreases the number of channels from 256 to 128 through a 1×1 convolutional layer (128 convolutional kernels, stride 1, padding=0). It focuses on the core information of "texture-semantic difference". For example, the medium texture difference (channel 60) and the deep semantic difference (channel 120) are merged into the feature of the new channel 60.

[0067] The fusion process employs "element-wise addition + ReLU activation." The two adapted 128-channel feature maps (both 128×128×128) undergo addition at each pixel location and corresponding channel. ReLU activation is then used to eliminate negative values, ensuring that the fused features only contain positive differences across scales. For example: Pixel (50,50), Channel 10: After shallow-to-medium layer fusion, the feature value is 0.2 (vegetation edge difference), after medium-to-deep layer fusion, the feature value is 0.3 (vegetation semantic difference), after addition it is 0.5, after ReLU it remains 0.5 (strong difference, corresponding to newly added vegetation); Pixels (80,80), Channel 60: After shallow-to-medium layer fusion, the feature value is 0.05 (no difference in road texture), after medium-to-deep layer fusion, the feature value is 0.03 (no difference in road semantics), after addition, it is 0.08, and after ReLU, it remains at 0.08 (weak difference, no change).

[0068] The optimized differential feature tensor is a 128×128×128 four-dimensional tensor (dimension order: height×width×number of channels). The channel dimension contains cross-scale difference information of "edge-texture-semantics": the first 40 channels focus on edge differences (such as changes in building outlines), channels 41-80 focus on texture differences (such as brick texture changing to vegetation texture), and channels 81-128 focus on semantic differences (such as "building-road" changing to "vegetation-road"). Regions with higher feature values ​​in the tensor have more significant cross-scale differences and a greater possibility of changes in ground features. For example, the feature values ​​of the original building area in the T2 image are mostly between 0.3 and 0.5, while the feature values ​​of the unchanged road area are mostly between 0.05 and 0.15, providing clear candidate regions of change for subsequent attention weighting.

[0069] S204, The spatial-channel dual-dimensional verification attention module is used to weight the differential features to generate dual-temporal attention enhancement features; Specifically, a spatial attention mechanism can be applied to the differential feature tensor, using a 7×7 convolution kernel to calculate the spatial location importance weights and generate a spatial attention weight map. This step focuses spatial attention on spatially varying regions with real changes in differentiated features, suppressing background noise and irrelevant areas. The core principle is to utilize large-size convolutional kernels to capture spatial correlation information, ensuring that weight allocation aligns with the spatial distribution characteristics of ground feature changes. The specific implementation is as follows: The differential feature tensor is a 128×128×128 three-dimensional structure (128 pixels high, 128 pixels wide, 128 channels). The channel dimension covers cross-scale differences between "edge, texture, and semantics". For example, the first 40 channels correspond to edge differences, channels 41-80 correspond to texture differences, and channels 81-128 correspond to semantic differences. The core logic of the spatial attention mechanism is to "model the spatial dependencies between pixels through convolution and assign high weights to spatially significant locations".

[0070] The reason for choosing a 7×7 convolutional kernel is that, compared to 3×3 or 5×5 convolutional kernels, the 7×7 convolution has a larger receptive field (covering 7×7=49 pixels), which can capture the spatial correlation of changes in ground features (such as the demolition area of ​​a building includes not only the building itself, but also the edge transition area of ​​1-2 pixels around it), avoiding misjudgment of "high response of isolated pixels" caused by an insufficiently small receptive field. The convolutional layer parameters are set as follows: 128 input channels, 1 output channel (compressing multi-channel features into single-channel spatial weights), stride 1, padding=3 (ensuring that the feature map size after convolution is still 128×128, consistent with the input), and the activation function is ReLU, suppressing negative weights (weights are 0 when spatial location is unimportant).

[0071] The calculation process consists of two steps: First, spatial information is integrated into the differential feature tensor through 7×7 convolution. For example, in the T2 image, the newly added vegetation region (pixels (50,50) to (70,70)) has high differential values ​​in its edge, texture, and semantic channels. After convolution, the single-channel feature value of this region reaches 0.8-0.9. The differential values ​​of the background road region (pixels (20,20) to (30,30)) are generally lower than 0.1, and the feature value after convolution is about 0.05-0.1. Then, the single-channel feature map output by convolution is normalized to map the feature values ​​to the 0-1 interval (normalization formula: weight=(value-min_value) / (max_value-min_value)), generating a spatial attention weight map.

[0072] The generated spatial attention weight map is a single-channel map of 128×128×1. The higher the weight value, the greater the probability of change in the corresponding spatial location. For example, the weight of newly added vegetation areas is 0.85-0.95, the weight of the edge area of ​​demolished buildings is 0.7-0.8, and the weight of unchanged roads and stable vegetation areas is 0.05-0.2. This clearly distinguishes between spatially changed and unchanged areas, providing a basis for the importance of spatial dimensions in subsequent weighting.

[0073] A channel attention mechanism is applied to the differential feature tensor. The importance score of each channel is calculated through global average pooling and fully connected layers to generate a channel attention weight vector. This step uses channel attention to filter the most effective channels for change detection among the differentiated features, suppressing redundant channels (such as channels that are sensitive to changes in illumination but unrelated to changes in ground features). The core is to use global information to evaluate the contribution of channels and ensure that the weight allocation fits the channel characteristics of the RGB image. The specific implementation is as follows: The starting point for calculating channel attention remains the 128×128×128 differential feature tensor. First, global average pooling (GAP) is performed, taking the arithmetic mean of all pixel values ​​for each channel. This compresses the two-dimensional channel features (128×128) into a one-dimensional scalar, with the formula: gap_c = (1 / (H×W)) × ΣΣfeature(i,j,c), where H=128 and W=128 are the feature map dimensions, and feature(i,j,c) is the feature value of pixel (i,j) in the c-th channel. For example: Channel 10 (focusing on architectural edge differences): all pixels have a mean of 0.35, gap_10=0.35; Channel 50 (focusing on vegetation texture differences): the mean value of all pixels is 0.42, gap_50=0.42; Channel 100 (focus-independent noise): mean value of all pixels is 0.08, gap_100=0.08; Global average pooling yields a 128-dimensional channel mean vector, with each element corresponding to the global importance basis of a channel.

[0074] Subsequently, a nonlinear transformation is performed on the channel mean vector through two fully connected layers (FC) to uncover the dependencies between channels: The first fully connected layer reduces the 128 dimensions to 32 dimensions (dimensionality reduction ratio of 4:1, balancing computational cost and feature representation), and the activation function is ReLU, with the formula fc1=ReLU(W1×gap+b1), where W1 is a 128×32 weight matrix and b1 is a 32-dimensional bias vector; The second fully connected layer increases the dimensionality back to 128 dimensions, and the activation function is Sigmoid, with the formula fc2=Sigmoid(W2×fc1+b2), where W2 is a 32×128 weight matrix and b2 is a 128-dimensional bias vector. The Sigmoid function ensures that the output channel weight values ​​are in the range of 0-1.

[0075] The generated channel attention weight vector is 128-dimensional, with each element corresponding to the importance score of a channel. A higher score indicates a greater contribution of that channel to change detection. For example, channels 41-60 (texture difference), which are related to vegetation change, score 0.7-0.9; channels 1-20 (edge ​​difference), which are related to building change, score 0.6-0.8; and channels 101-128, which are related to noise, score 0.1-0.3. This vector accurately selects the change-sensitive channels in the RGB image, avoiding redundant channels from interfering with subsequent feature weighting.

[0076] The spatial attention weight map and the channel attention weight vector are broadcast and added together, and a preliminary fused attention map is generated by the Sigmoid function; This step resolves the dimensionality difference between the spatial weight map and the channel weight vector through a broadcast mechanism, achieving collaborative fusion of two-dimensional attention. The core is to ensure that each pixel-channel position receives a comprehensive weight. The specific implementation is as follows: First, clarify the dimensional difference between the two: the spatial attention weight map is 128×128×1 (H×W×1), and the channel attention weight vector is 1×1×128 (1×1×C). Directly adding them results in a dimensional mismatch problem. Therefore, it is necessary to expand the dimensions to a consistent 128×128×128 (H×W×C) through a "broadcast" operation. Spatial weight map broadcasting: The 128×128×1 weight map is copied 128 times along the channel dimension to obtain a 128×128×128 spatial weight tensor. The spatial weight distribution of each channel is completely consistent (for example, the spatial weight maps of the 10th and 50th channels are copies of the original 128×128×1). Channel weight vector broadcasting: The 1×1×128 weight vector is copied 128 times along the height and width dimensions to obtain a 128×128×128 channel weight tensor. The channel weight distribution at each spatial location is completely consistent (for example, the channel weight vectors of pixel (50,50) and pixel (80,80) are both copies of the original 1×1×128).

[0077] After broadcasting, the spatial weight tensor and the channel weight tensor are added element-wise to obtain the fused weight tensor. The addition formula is: fusion_weight(i,j,c) = spatial_weight(i,j,c) + channel_weight(i,j,c). For example: Pixel (50,50), 50th channel (vegetation texture + spatial high weight): spatial_weight=0.85, channel_weight=0.8, fusion_weight=1.65; Pixel (20,20), 100th channel (noise + low spatial weight): spatial_weight=0.1, channel_weight=0.2, fusion_weight=0.3; After addition, the fusion weight tensor is normalized using the Sigmoid function, mapping all elements to the 0-1 range. The formula is: pre_attention(i,j,c)=Sigmoid(fusion_weight(i,j,c)). The characteristic of the Sigmoid function is that it outputs a value close to 1 for high fusion weights (such as 1.65) and a value close to 0 for low fusion weights (such as 0.3).

[0078] The generated preliminary fusion attention map is a 128×128×128 three-dimensional tensor, where each element represents the comprehensive importance weight of the corresponding "pixel-channel" position. For example, the weight of "pixel (55,55) - channel 50" in the newly added vegetation area is 0.92, the weight of "pixel (70,70) - channel 15" in the demolished building area is 0.88, and the weight of "pixel (25,25) - channel 110" in the unchanged road area is 0.15, thus initially achieving the dual enhancement of "spatial focusing + channel filtering".

[0079] By combining the feature alignment verification mechanism, the cosine similarity of the attention features at time T1 and T2 is calculated. Weight penalties are applied to similarity regions that exceed the preset similarity threshold, a refined attention map is generated and weighted to the differentiated features, and finally the dual-temporal attention enhancement features are output.

[0080] This step eliminates pseudo-change regions (such as feature differences caused by changes in illumination) through feature alignment verification. The core is to use the similarity of bi-temporal features to determine the authenticity of changes, ensuring that the attention weights focus only on genuine changes. The specific implementation is as follows: First, obtain the attention features at times T1 and T2: multiply the preliminary fused attention map element-wise with the differential feature tensors at times T1 and T2 respectively to obtain the T1 attention feature tensor (att_T1=pre_attention×feat_T1) and the T2 attention feature tensor (att_T2=pre_attention×feat_T2), both of which have a dimension of 128×128×128, preserving the "attention-feature" association of each time phase.

[0081] Calculate cosine similarity: Calculate the cosine similarity cos_sim(i,j,c) for the corresponding pixel-channel positions of att_T1 and att_T2 to measure the degree of similarity between the two temporal features.

[0082] Preset similarity threshold: Based on a large amount of RGB dual-temporal data statistics, the cosine similarity of pseudo-change regions is generally higher than 0.8, while that of real change regions is mostly lower than 0.5. Therefore, the threshold is set to 0.8. Weight penalties are applied to regions where cos_sim(i,j,c) > 0.8. The penalty formula is: refine_weight(i,j,c) = pre_attention(i,j,c) × (1 - (cos_sim(i,j,c) - 0.8) / 0.2), where (1 - ...) is the penalty coefficient. When cos_sim = 0.8, the penalty coefficient = 1 (no penalty); when cos_sim = 1.0, the penalty coefficient = 0 (full penalty). For example: cos_sim=0.9 (strong pseudo-change): refine_weight=0.8×(1-(0.9-0.8) / 0.2)=0.8×0.5=0.4 (weight halved); cos_sim=0.7 (actual change): refine_weight=0.92×1=0.92 (no penalty).

[0083] After generating the refined attention map, it is multiplied element-wise with the differential feature tensor to obtain the bi-temporal attention enhancement feature: enhanced_feat = refined_weight × diff_feat. For example: The actual change area (pixels (55,55) - 50th channel): refine_weight=0.92, diff_feat=0.5, enhanced_feat=0.46; False change region (pixels (30,30) - 80th channel): refine_weight=0.4, diff_feat=0.3, enhanced_feat=0.12; The final output of the dual-temporal attention-enhanced feature is a 128×128×128 tensor. The feature values ​​of the real change regions are significantly enhanced, while the pseudo-change regions are suppressed, providing high-purity change feature input for subsequent difference calculation and classification.

[0084] S205, perform difference calculation and classification on the dual-temporal attention enhancement features, and output the land cover change area.

[0085] Specifically, pixel-wise absolute value difference operations can be performed on the dual-temporal attention enhancement features to obtain an initial feature difference map; This step transforms abstract attention-enhancing features into intuitive difference representations through direct comparison of bi-temporal features. The core principle is to leverage absolute value differences to highlight the true changes across both temporal phases, while avoiding information loss caused by the cancellation of positive and negative differences. The specific implementation is as follows: The dual-temporal attention enhancement features include a feature tensor at time T1 (att_feat_T1) and a feature tensor at time T2 (att_feat_T2), both of which have a dimension of 128×128×128 (height 128 pixels, width 128 pixels, channels 128). The channel dimension covers attention-weighted features of "edge-texture-semantics". For example, channel 10 corresponds to building edge enhancement features, channel 50 corresponds to vegetation texture enhancement features, and channel 100 corresponds to ground feature semantic enhancement features.

[0086] The pixel-by-pixel absolute difference operation is performed on the pixel-to-channel positions of two temporal features one by one. The formula is: diff(i,j,c)=|att_feat_T2(i,j,c)-att_feat_T1(i,j,c)|, where (i,j) are the pixel coordinates and c is the channel index. The absolute value operation ensures that the differences between "T2 feature is higher than T1" and "T1 feature is higher than T2" are preserved (the former corresponds to newly added features, and the latter corresponds to disappeared features). For example: New vegetation area (pixel (50,50), 50th channel): att_feat_T1=0.05 (T1 has no vegetation, low feature after attention weighting), att_feat_T2=0.85 (T2 has vegetation, high feature after attention weighting), diff=|0.85-0.05|=0.8 (high difference, corresponding to new changes); Demolition area (pixels (70,70), 10th channel): att_feat_T1=0.9 (T1 has buildings, high edge features), att_feat_T2=0.1 (T2 has buildings demolished, low edge features), diff=|0.1-0.9|=0.8 (high difference, corresponding to vanishing changes); Unchanged road area (pixels (20,20), 80th channel): att_feat_T1=0.2, att_feat_T2=0.22, diff=|0.22-0.2|=0.02 (low difference, no change).

[0087] After computation, the 128-channel difference features need to be compressed into a single-channel initial feature difference map. Channel max pooling (taking the maximum difference value for each pixel along the channel dimension) is used, with the formula: `init_diff_map(i,j)=max_c(diff(i,j,c))`. This ensures that the difference value of each pixel reflects its most significant change across all channels. For example, if the maximum diff for pixel (50,50) in the 128 channels is 0.8 (channel 50), then `init_diff_map(50,50)=0.8`; if the maximum diff for pixel (20,20) is 0.03 (channel 80), then `init_diff_map(20,20)=0.03`.

[0088] The generated initial feature difference map is a single-channel image of 128×128×1 with pixel values ​​ranging from 0 to 1. High-value areas (0.5-1.0) correspond to potential changes, while low-value areas (0-0.2) correspond to no changes. It has initially presented the spatial distribution of land cover changes, but there are problems such as local noise (e.g., high differences in single pixels) and small patch breaks (e.g., only some pixels of small vegetation patches ≤10 pixels have high response), which require further processing and optimization.

[0089] Spatial context features are extracted from the initial feature difference map using a 3×3 convolutional layer to enhance the continuous representation of the changing regions and generate a context-enhanced difference map. This step eliminates local noise and connects broken and changing regions through spatial context integration. The core is to use small-sized convolutional kernels to capture local dependencies between pixels, ensuring the spatial continuity of changing regions. This is suitable for detecting small patches of ground features (such as scattered vegetation and small buildings) in RGB images. The specific implementation is as follows: The initial feature difference map is a single-channel image of 128×128×1, which suffers from the problems of "isolated high-response pixels" (such as single pixels with a diff=0.6 caused by road stains) and "fragmented small patches" (such as small vegetation patches of 5×5 pixels with only 3 pixels and a diff>0.5). The role of the 3×3 convolutional layer is to integrate isolated high-response pixels into surrounding low-response pixels through local pixel weighted fusion (suppressing noise) and connect fragmented high-response pixels into continuous regions (enhancing small patches).

[0090] The convolutional layer parameters are set as follows: input channel 1, output channel 1 (preserving the single-channel difference map property), stride 1, padding=1 (edge ​​padding ensures the image size remains 128×128 after convolution, avoiding loss of edge pixel information). The convolutional kernel weights adopt a Gaussian distribution design with "high center, low edge", for example, the weight matrix is: [[0.05,0.1,0.05],[0.1,0.5,0.1],[0.05,0.1,0.05]]. This weight makes the convolution result focus more on the difference value of the current pixel, while taking into account the contextual information of the surrounding 8 pixels, balancing "local realism" and "regional continuity".

[0091] Example of calculation process: Isolated noisy pixel (20,20): init_diff_map(20,20)=0.6, the diff of the surrounding 8 pixels is 0.02-0.05, the result after convolution = 0.6×0.5+0.05×0.1×4+0.02×0.05×4=0.3+0.02+0.004=0.324 (the noise is suppressed, and the difference value decreases from 0.6 to 0.324); For the broken small patch (50,50) to (54,54): the center 3 pixels have a diff=0.8, the surrounding 2 pixels have a diff=0.3, and the remaining pixels have a diff=0.05. After convolution, the center pixel result is 0.8×0.5+0.8×0.1×2+0.3×0.1×2+0.05×0.05×4=0.4+0.16+0.06+0.01=0.63 (the low response of the surrounding area is increased, and the small patch is made continuous).

[0092] The generated context-enhanced difference map is still 128×128×1 with pixel values ​​ranging from 0 to 0.9. Compared with the initial difference map, the proportion of isolated noise has decreased from 15% to 5%, and the continuity rate of small patches with a resolution of ≤10 pixels has increased from 40% to 85%. For example, the originally broken 5×5 vegetation patches are transformed into complete 5×5 high-difference regions (diff>0.5) after convolution, providing a more reliable representation of the change region for subsequent classification.

[0093] The context-enhanced difference map is input into the dual-branch classifier, and the pixel-level change probability is calculated through a fully connected layer to generate a preliminary change probability map; This step transforms the difference map into a probabilistic representation using a dual-branch classifier. The core principle is to capture change features (edge ​​texture and semantic association) across dimensions, ensuring the accuracy of pixel-level change judgments and avoiding misjudgments caused by single-dimensional features. The specific implementation is as follows: The dual-branch classifier architecture is adapted to the feature characteristics of RGB images, consisting of an "edge texture branch" and a "semantic association branch." The two branches share the input (context-enhanced difference map), but emphasize different variation features: Edge texture branch: The single-channel difference map is mapped to 32-dimensional features through 1×1 convolution (focusing on local changes in edges and textures), and then outputs the probability of local changes through two fully connected layers (the first layer has 32→16 neurons, and the second layer has 16→1 neurons). Semantic association branch: Extract land cover association features (focusing on semantic changes of “vegetation-bare land” and “building-road”) through 3×3 convolution (stride 1, padding=1), and map them to 32-dimensional features. Output the global association change probability through two fully connected layers.

[0094] The activation function for all fully connected layers is Sigmoid, mapping the output values ​​to the 0-1 range, where 0 represents "100% probability of no change" and 1 represents "100% probability of change". The classifier is trained on a labeled RGB biphase dataset (such as LEVIR-CD), using binary cross-entropy as the loss function and Adam as the optimizer (learning rate 0.001) to ensure the model learns the mapping relationship between "diff value magnitude → probability of change". For example, diff=0.8 corresponds to a probability of 0.9, and diff=0.05 corresponds to a probability of 0.05.

[0095] Example of calculation process: New vegetation area (pixel (50,50)): Context enhancement difference value = 0.63, edge texture branch output probability 0.88 (significant texture difference), semantic association branch output probability 0.92 ("bare land → vegetation" semantic change), the final pixel-level change probability is the average of the two branches: (0.88+0.92) / 2=0.9; Unchanged road region (pixels (20,20)): Context-enhanced difference value = 0.03, edge texture branch output probability 0.04, semantic association branch output probability 0.02, mean = 0.03; Small patch vegetation region (pixels (60,60)): Context-enhanced difference value = 0.55, edge texture branch output probability 0.8, semantic association branch output probability 0.75, mean = 0.775.

[0096] The generated preliminary change probability map is a single-channel image of 128×128×1, with pixel values ​​ranging from 0 to 1. High probability areas (0.7-1.0) correspond to definite changes, medium probability areas (0.3-0.7) correspond to suspected changes, and low probability areas (0-0.3) correspond to no changes. The probability of change for each pixel has been quantified, but there are still problems of "misjudgment in medium probability areas" and "isolated high probability noise points", which need further optimization.

[0097] Adaptive threshold segmentation is applied to the preliminary change probability map, and morphological post-processing is used to eliminate isolated noise points, ultimately outputting an accurate binary map of land cover change areas.

[0098] This step transforms the probability map into a precise binary variation region through threshold segmentation and morphological operations. The core is adaptive thresholding to fit the probability distribution of different regions. Morphological post-processing eliminates noise and fills in holes, ensuring the output reflects the spatial morphology of actual ground features. The specific implementation is as follows: Adaptive threshold segmentation The Otsu algorithm (maximum inter-class variance method) is used to determine the segmentation threshold, avoiding "false positives in bright areas" or "false negatives in dark areas" caused by a fixed threshold (such as 0.5). The core logic of the Otsu algorithm is: iterate through all possible thresholds T in the 0-1 interval, calculate the inter-class variance of the probability map divided into "foreground (changing)" and "background (unchanging)" by threshold T, and take the T with the largest inter-class variance as the optimal threshold. For example: The probability distribution of the preliminary change probability map is as follows: the area with no change (0-0.3) accounts for 70%, the area with suspected change (0.3-0.7) accounts for 20%, and the area with clear change (0.7-1.0) accounts for 10%. The Otsu algorithm calculates that the inter-class variance is maximum when T=0.65. That is, pixels with a probability ≥0.65 are judged as "change" (marked as 1), and pixels with a probability <0.65 are judged as "no change" (marked as 0). At this time, the separation between the foreground and the background is the highest, 70% of the real changes in the medium probability region are correctly identified, and the misclassification rate is reduced to below 5%.

[0099] After segmentation, a preliminary binarized image is generated, but there are still two types of problems: one is isolated noise points (such as a single pixel with a probability of 0.7 and surrounding pixels with a probability of 0.1, corresponding to pseudo-changes); the other is holes in the change region (such as a 2×2 pixel in the center of a 10×10 change region with a probability of 0.6, which is judged as no change and forms a hole).

[0100] Morphological post-processing Use a combination of "opening operation" and "closing operation": Opening operation (erosion followed by dilation): eliminates isolated noise points. The erosion operation uses a 3×3 square structuring element (all pixels within the structuring element are 1), traversing the binarized image. Only when all 9 pixels covered by the structuring element are 1 is the center pixel retained as 1; otherwise, it becomes 0. For example, a single-pixel noise point becomes 0 after erosion. The dilation operation uses the same structuring element, traversing the eroded image. If even one pixel covered by the structuring element is 1, the center pixel becomes 1, ensuring the size of the effective change area does not shrink. For example: an isolated noise point (1 pixel 1) becomes 0 after erosion, then 0 after dilation (noise elimination); a 5×5 change area becomes 3×3 after erosion, then 5×5 after dilation (size restoration).

[0101] Closing operation (dilation followed by erosion): fills holes in a changing region. The dilation operation first expands the 1-pixel edge of the hole into the hole itself; for example, a 2×2 hole becomes 1 pixel after dilation. The erosion operation then shrinks the excess expanded pixels, ensuring that the shape of the changing region remains unchanged. For example, a 2×2 hole in the center of a 10×10 changing region disappears after dilation, and then the region shape remains 10×10 (hole filling) after erosion.

[0102] After morphological operations, a final binarized map of changes in land cover is generated, which is a single-channel image of 128×128×1 pixels. The pixel values ​​are only 0 or 1, where 1 represents a changed area and 0 represents a unchanged area. For example, newly added vegetation areas are fully marked as 1 (a continuous block of 50×50 pixels), demolished building areas are marked as 1 (a continuous block of 30×30 pixels), and roads and stable vegetation areas with no change are marked as 0. Isolated noise points and internal cavities are eliminated, and the edges of the changed areas are smooth and the morphology is complete, perfectly matching the spatial distribution of actual land cover changes.

[0103] The binarized image can be further mapped back to the pixel coordinates of the original RGB image (by inverse mapping of size, scaling from 128×128 to the original 512×512 or 1024×1024), and the output is a visualized change detection result, such as marking the change area with a red outline on the original T2 image, which is convenient for manual verification and subsequent applications (such as illegal construction monitoring and disaster assessment).

[0104] As can be seen, radiometric consistency correction and joint data augmentation are performed on the dual-temporal RGB images to obtain preprocessed dual-temporal images. A weight-sharing twin backbone network is used to extract multi-scale features from the preprocessed dual-temporal images. Redundant information in the multi-scale features is ablated using a multi-scale feature complementation module to obtain differentiated features. These differentiated features are weighted using a spatial-channel dual-dimensional verification attention module to generate dual-temporal attention-enhanced features. Difference calculation and classification are performed on the dual-temporal attention-enhanced features to output the change areas of ground features. This improves the accuracy and robustness of change detection, overcomes environmental noise interference, and achieves more complete and accurate identification of multi-scale ground feature changes.

[0105] Another embodiment of the present invention provides a ground cover change detection system based on multi-scale feature fusion and attention mechanism, see [link to relevant documentation]. Figure 3 The system may include: Processing module 301 is used to perform radiometric consistency correction and joint data enhancement on the dual-temporal RGB image to obtain the preprocessed dual-temporal image; Extraction module 302 is used to extract multi-scale features of the preprocessed dual-temporal image using a weight-sharing twin backbone network; The ablation module 303 is used to ablate redundant information of the multi-scale features through the multi-scale feature complementarity module to obtain differentiated features; Weighting module 304 is used to weight the differential features using the spatial-channel dual-dimensional verification attention module to generate dual-temporal attention enhancement features; The output module 305 is used to perform difference calculation and classification on the dual-temporal attention enhancement features and output the land cover change area.

[0106] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0107] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0108] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.

[0109] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A method for feature change detection of ground objects based on multi-scale feature fusion and attention mechanism, characterized in that, The method comprises: radiometric consistency correction and joint data enhancement are performed on the dual-time RGB images to obtain preprocessed dual-time images; a weight-shared twin backbone network is used to extract multi-scale features of the preprocessed dual-time images; redundant information of the multi-scale features is eliminated through a multi-scale feature complementary module to obtain differentiated features; a spatial-channel two-dimensional verification attention module is used to weight the differentiated features to generate dual-time attention enhanced features; difference calculation and classification are performed on the dual-time attention enhanced features to output the feature change area.

2. The method of claim 1, wherein, The method comprises: original RGB images at T1 and T2 are loaded respectively, image data is converted into a floating-point tensor format, and original dual-time image tensors are generated; radiometric consistency correction is performed on the original dual-time image tensors, linear regression algorithm is used to calculate the mean and variance of each channel of the T1 and T2 images, and the channel values of the T2 image are adjusted to align with the statistical characteristics of the T1 image to generate radiometric correction image tensors; joint data enhancement is performed on the radiometric correction image tensors, random rotation, horizontal flipping and scale zooming operations are sequentially performed, and color jittering is applied to simulate illumination changes to generate enhanced image tensors; the enhanced image tensors are uniformly scaled to a specific pixel size, and channel-level standardization processing is performed to finally generate preprocessed dual-time images.

3. The method of claim 2, wherein, The method comprises: a weight-shared twin backbone network is constructed, four feature extraction stages are designed based on a lightweight ResNeSt50 architecture, and a twin network model is generated; the preprocessed dual-time images are input into two branches of the twin network respectively, pixel-level edge features and color distribution features are extracted through the first stage convolutional layer to generate shallow feature maps; local texture features are extracted through the second stage convolutional layer combined with the maximum pooling operation, and grouping convolution is used to enhance feature diversity to generate middle-layer texture feature maps; global semantic features are extracted through the third stage convolutional layer by introducing dilated convolution with dilation rates of 2, 4 and 6 to expand the receptive field, and finally a feature pyramid containing multi-scale information is generated.

4. The method of claim 3, wherein, The method comprises: feature maps of adjacent levels are selected from the multi-scale feature pyramid, the feature maps include shallow-middle and middle-deep feature combinations to generate a feature map pair set; element-wise subtraction is performed on each feature map pair to calculate a difference matrix between adjacent scale features, eliminate repetitive texture patterns, and generate initial difference features; local response normalization is performed on the initial difference features to suppress background noise and enhance the response intensity of significant change areas to generate enhanced difference features; all enhanced difference features are fused through a 1×1 convolutional layer to retain differentiated information across scales, and finally an optimized differentiated feature tensor is generated.

5. The method of claim 4, wherein, The space-channel dual-dimension verification attention module is used for weighting the differential feature, to generate a dual-time-phase attention enhanced feature, including: A spatial attention mechanism is applied to the differential feature tensor, a 7*7 convolution kernel is used to calculate the spatial position importance weight, and a spatial attention weight map is generated; A channel attention mechanism is applied to the differential feature tensor, a global average pooling and a fully connected layer are used to calculate the importance score of each channel, and a channel attention weight vector is generated; The spatial attention weight map and the channel attention weight vector are added by broadcasting, and a preliminary fusion attention map is generated by a Sigmoid function; Combined with the feature alignment verification mechanism, the cosine similarity of the T1 and T2 time attention features is calculated, the similarity regions higher than the preset similarity threshold are weighted and punished, a refined attention map is generated, and is weighted to the differential feature, and finally a dual-time-phase attention enhanced feature is output.

6. The method of claim 5, wherein, The dual-time-phase attention enhanced feature is subjected to difference calculation and classification, and a ground feature change region is output, including: A pixel-by-pixel absolute value difference operation is performed on the dual-time-phase attention enhanced feature to obtain an initial feature difference map; A 3*3 convolution layer is used to extract spatial context features from the initial feature difference map, to enhance the continuity representation of the change region, and generate a context enhanced difference map; The context enhanced difference map is input into a dual-branch classifier, a fully connected layer is used to calculate the pixel-level change probability, and a preliminary change probability map is generated; An adaptive threshold segmentation is performed on the preliminary change probability map, and isolated noise points are removed by morphological post-processing, to finally output an accurate binary ground feature change region map.

7. A land feature change detection system based on multi-scale feature fusion and attention mechanism, characterized in that, The system comprises: A processing module configured to perform radiation consistency correction and joint data enhancement on dual-time-phase RGB images to obtain preprocessed dual-time-phase images; An extraction module configured to extract multi-scale features of the preprocessed dual-time-phase images by using a weight-shared twin backbone network; An ablation module configured to ablate redundant information of the multi-scale features by using a multi-scale feature complementary module to obtain differential features; A weighting module configured to weight the differential features by using a space-channel dual-dimension verification attention module to generate dual-time-phase attention enhanced features; An output module configured to perform difference calculation and classification on the dual-time-phase attention enhanced features to output a ground feature change region.

8. The system of claim 7, wherein, The processing module is specifically configured to: Load original RGB images at T1 and T2 time respectively, convert the image data into a floating-point tensor format, and generate original dual-time-phase image tensors; Perform radiation consistency correction on the original dual-time-phase image tensors, calculate the mean and variance of each channel of the T1 and T2 images by using a linear regression algorithm, and adjust the channel values of the T2 image to align with the statistical features of the T1 image to generate radiation corrected image tensors; Perform joint data enhancement on the radiation corrected image tensors, sequentially perform random rotation, horizontal flip and scale operation, and simultaneously apply color jittering to simulate illumination changes to generate enhanced image tensors; Uniformly scale the enhanced image tensors to a specific pixel size, and perform channel-level standardization processing to finally generate preprocessed dual-time-phase images.

9. A storage medium, characterized by The storage medium has stored therein a computer program, wherein the computer program is configured to execute the method of any one of claims 1-6 when executed.

10. An electronic device comprising a memory and a processor, characterized in that, The memory has stored therein a computer program, and the processor is configured to execute the computer program to execute the method of any one of claims 1-6.