Remote sensing image change detection method based on depth adaptive attention mechanism
Through the deep adaptive attention mechanism network model and Gaussian splash optimization technology, the accuracy and robustness problems in complex scenes in remote sensing image change detection are solved, and efficient and accurate identification of change areas is achieved.
Patent Information
- Application Number
- CN202510476554.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has insufficient change detection accuracy, multimodal data fusion capability and dynamic modeling capabilities in complex scenarios, making it difficult to meet the high accuracy and high robustness requirements of multi-source remote sensing image change detection tasks.
The remote sensing image change detection method based on the depth adaptive attention mechanism is adopted, and feature extraction and multimodal fusion are performed through the deep adaptive attention mechanism network model, combined with Gaussian splash optimization and attention allocation, the weight is dynamically adjusted, and a fusion feature map is generated to identify the changing area.
It significantly improves the efficiency and robustness of remote sensing image change detection, enhances the ability to identify key changing areas, reduces the missed detection and error detection rates, and improves the detection accuracy of fine-grained change areas.
Smart Images

Figure CN120375073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing images, and in particular, to a remote sensing image change detection method based on a deep adaptive attention mechanism. Background Art
[0002] With the development of deep learning and remote sensing technology, change detection of remote sensing images has been widely applied in fields such as environmental monitoring, urban planning, agricultural assessment, and disaster response. The core of change detection lies in extracting change regions and their features from multi-temporal remote sensing images, so as to achieve the monitoring of dynamic changes in the target area.
[0003] Currently, traditional change detection methods mainly include change extraction techniques based on pixel difference, classifier training, and image registration. These techniques usually rely on relatively simple statistical models or rule designs, and have poor adaptability to noise, shadow interference, and complex backgrounds in image data. In cases of vegetation coverage, dense urban building areas, or high image noise, traditional methods are prone to missed detections or false detections. In addition, they have insufficient capabilities in multi-source data fusion, usually only being able to process single-modal data and unable to fully utilize the complementarity of data such as optical images, multi-spectral images, and synthetic aperture radar images.
[0004] In recent years, change detection methods based on deep learning have gradually attracted attention. Deep learning extracts deep features and identifies changes in remote sensing images through neural network models, improving the accuracy of change detection. However, existing deep learning methods usually rely on fixed network architectures and feature allocation strategies, and are unable to dynamically adjust the computational focus of the network according to the differences in image content. There are still significant deficiencies in the modeling capabilities of key regions and the accuracy of feature fusion when processing high-resolution images and multi-modal data. In addition, most deep learning methods fail to effectively utilize the spatial context information and local detail features in remote sensing images, and there are certain difficulties in detecting fine-grained change regions.
[0005] In summary, the existing technologies have significant deficiencies in the change detection accuracy in complex scenarios, the capabilities of multi-modal data fusion, and the dynamic modeling capabilities of key regions, and it is difficult to meet the high-precision and high-robustness requirements of multi-source remote sensing image change detection tasks. Therefore, there is an urgent need for a new change detection method to solve the above problems. Summary of the Invention
[0006] An object of the present invention is to propose a remote sensing image change detection method based on a deep adaptive attention mechanism, which significantly improves the efficiency and robustness of the remote sensing image change detection task.
[0007] A remote sensing image change detection method based on a deep adaptive attention mechanism according to an embodiment of the present invention includes the following steps:
[0008] S1. Obtain multi-temporal remote sensing image data, where the remote sensing image data includes optical images, multi-spectral images, and synthetic aperture radar images. Preprocess the remote sensing image data through radiometric calibration, geometric calibration, and image registration to obtain preprocessed remote sensing image data that is spatially aligned and spectrally consistent;
[0009] S2. Input the preprocessed remote sensing image data into a deep adaptive attention mechanism network model. The deep adaptive attention mechanism network model includes a feature extraction module, an attention allocation module, and a multi-modal fusion module. Use the feature extraction module to perform multi-scale feature extraction on the preprocessed remote sensing image data to generate a preliminary multi-scale feature map;
[0010] S3. Perform Gaussian splash processing on the preliminary multi-scale feature map. Based on the change center point and neighborhood distribution, diffuse the features through Gaussian distribution weights to generate a final Gaussian splash optimized feature map;
[0011] S4. Use the attention allocation module to perform dynamic weight allocation on the Gaussian splash optimized feature map, adjust the weight distribution according to the importance of the feature regions, and generate an optimized weighted feature map. The optimized weighted feature map retains the significant features of the key change regions;
[0012] S5. Input the optimized weighted feature map into the multi-modal fusion module, combine the feature information of the optical image, multi-spectral image, and synthetic aperture radar image to generate a fusion feature map. The fusion feature map represents the change information in the multi-source remote sensing image data in the form of a high-dimensional feature space;
[0013] S6. Input the fusion feature map into the change detection module to perform classification processing on the fusion feature map, identify the change regions, and generate a change detection result, including the spatial distribution of the change regions and the annotation information of the change types;
[0014] S7. Post-process the change detection result, use morphological operations to optimize the boundary continuity of the change regions and the integrity of the detection result, and finally generate an annotated remote sensing image of the change regions.
[0015] Optionally, the S1 includes:
[0016] S11. Obtain multi-temporal remote sensing image data, where the remote sensing image data includes optical image I optical , multi-spectral image I optical and synthetic aperture radar image I SAR ;
[0017] S12. Perform radiometric calibration on the remote sensing image data;
[0018] S13. Geometrically correct the radiometrically corrected remote sensing image data by constructing a mapping function between the image and geographic coordinates. The mapping function is described using a polynomial transformation model:
[0019]
[0020] where T(x, y) represents the projection point of the image coordinate point in the geographic coordinate system, x and y represent the image coordinates respectively, a ij is the model coefficient of geometric correction, and n is the order of the polynomial;
[0021] S14. Perform image registration on the geometrically corrected remote sensing image data. Adopt the method of feature point matching. By selecting the reference image and the image to be registered, extract the image feature point set, and use affine transformation for spatial alignment;
[0022] S15. Combine and process the remote sensing image data that has undergone radiometric correction, geometric correction, and image registration to obtain preprocessed remote sensing image data with spatial alignment and spectral consistency:
[0023] I preprocessed ={I optical,pre , I multispectral,pre , I SAR,pre};
[0024] where I optical,pre , I multispectral,pre and I SAR,pre represent the preprocessed optical image, multispectral image, and synthetic aperture radar image respectively.
[0025] Optionally, the above S2 includes:
[0026] S21. Input the preprocessed remote sensing image data I preprocessed into the deep adaptive attention mechanism network model, and the deep adaptive attention mechanism network model includes a feature extraction module, an attention allocation module, and a multimodal fusion module;
[0027] S22. Perform embedding mapping on different types of preprocessed remote sensing image data respectively in the feature extraction module to obtain an initial multi-scale feature representation:
[0028]
[0029] where I m,pre represents the input data of the m-th type of image, m is the image type index, including optical, multispectral, and synthetic aperture radar, l represents the multi-scale level, is the feature mapping function of the l-th layer;
[0030] S23. Input the initial multi-scale feature representation into the attention allocation module, and use the multi-head self-attention mechanism to perform deep recalibration on the features of each type of image:
[0031]
[0032] Among them, respectively represent the query, key, and value vectors for the m-th type of image feature in the l-th layer, are the corresponding trainable weight parameters;
[0033] S24. Perform weighted operations on the query, key, and value vectors in the attention allocation module to generate a deep adaptive attention feature representation for the multi-scale features of each type of image:
[0034]
[0035] Among them, represents the attention feature that fuses local and global context information, and Attn(·) is the multi-head self-attention function, which strengthens the key regions of multi-modal data;
[0036] S25. Use the multi-modal fusion module to process the deep adaptive attention features obtained from different types of images:
[0037]
[0038] Perform cross-modal aggregation to generate a preliminary multi-scale feature map:
[0039]
[0040] Among them, is a learnable weighting coefficient used to coordinate the importance of features of different image types;
[0041] S26. Output a preliminary multi-scale feature map containing multiple scale levels:
[0042]
[0043] Among them, L is the total number of multi-scale levels.
[0044] Optionally, the S3 includes:
[0045] S31. Input the preliminary multi-scale feature map F combined into the Gaussian splash processing module to determine the set P center ={p1, p2,..., p n} of the initial feature distribution center points of the changing regions, where p i is the coordinate of the i-th changing region center point, and n is the number of changing center points;
[0046] S32. According to the set of change center points P center Perform weight assignment on the feature image pixels around each center point. The weight is defined by a two-dimensional Gaussian distribution function as:
[0047]
[0048] where G(x, y; μ x , μ y , σ x , σ y ) is the Gaussian distribution weight, representing the weight value at position (x, y), (μ x , μ y ) represents the coordinates of the change center point, and σ x and σ y represent the horizontal and vertical standard deviations of Gaussian diffusion respectively;
[0049] S33. For the preliminary feature map at each scale l According to the Gaussian distribution weight G(x, y; μ x , μ y , σ x , σ y ) perform weighted calculation on the eigenvalue to generate a Gaussian splash optimized feature map
[0050]
[0051] where is the optimized eigenvalue at position (x, y), enhancing the feature expression in the change region and suppressing the redundant information in the background region;
[0052] S34. At each scale l, combine the weight superposition feature distributions of all change center points to calculate the final Gaussian splash optimized feature map:
[0053]
[0054] where represents the optimized feature component generated by the i-th change center point;
[0055] S35. Combine the Gaussian splash optimized feature maps at each scale l and output the set of final Gaussian splash optimized feature maps:
[0056]
[0057] Optionally, the S4 includes:
[0058] S41. Input the set of feature maps optimized by Gaussian splashing into the attention allocation module, initialize the query, key, and value vectors, and define the query vector Q (l) , the key vector K (l) and the value vector K (l) are respectively:
[0059]
[0060] where W Q(l) , W K(l) and W V(l) are the trainable weight matrices of the l-th layer, used to generate the query, key, and value vectors;
[0061] S42. Calculate the attention weight matrix A (l) for the query vector Q (l) and the key vector K (l) ;
[0062] S43. Perform weighted summation on the value vector V (l) according to the attention weight matrix A (l) to generate the optimized attention feature map
[0063]
[0064] where, represents the optimized feature map that retains the significant features in the key change regions;
[0065] S44. Introduce a residual connection to the optimized attention feature map to enhance feature transmission:
[0066]
[0067] where, is the final feature map that fuses the attention mechanism and Gaussian splashing optimization information;
[0068] S45. Aggregate the final feature maps of all scale levels to generate the optimized set of weight feature maps:
[0069]
[0070] Optionally, the S5 includes:
[0071] S51. Input the optimized set of weight feature maps into the multi-modal fusion module, and perform modal weighting on the features of the optical image, multi-spectral image, and synthetic aperture radar image respectively:
[0072]
[0073] where, Denote the fusion feature of the m-th modality at the position (x, y). is a learnable modality weighting coefficient used to adjust the contribution of each modality to feature fusion. is a gradient smoothing factor used to suppress noise in the features. is the sum of the squares of the feature gradients, representing the change intensity near the position (x, y).
[0074] S53. Perform cross-modal superposition on different modality features to generate a fused feature map:
[0075]
[0076] Among them, represents the value of the fused feature map at the position (x, y), w m is the superposition weight of the modality feature, and ReLU(·) is an activation function used to strengthen the expression of positive features. is a non-linear regularization term.
[0077] S54. Introduce spatial attention weights to dynamically adjust the local weights of the fused feature map and calculate the attention weight matrix:
[0078]
[0079] Among them, S (l) (x, y) is the spatial attention weight of the fused feature map at the position (x, y). represents the adaptive scaling factor of the fused feature.
[0080] S55. Optimize the fused feature map based on the spatial attention weights to generate the final multi-modal fused feature map:
[0081]
[0082] Among them, is the final fused feature at the position (x, y). is the activation strengthening term.
[0083] The beneficial effects of the present invention are:
[0084] (1) The present invention adopts a deep adaptive attention mechanism. By introducing a dynamic weight adjustment strategy in the feature extraction and allocation process, the model can focus on key change regions and suppress the interference of redundant information. The dynamic attention allocation mechanism can automatically adjust the attention weights according to the image content, strengthening the precise recognition ability of change regions. In scenes with high vegetation coverage, complex building shadows, and high noise, through the Gaussian splash optimization of the preliminary multi-scale feature maps, the feature expression of local change regions is enhanced, solving the problems of missed detection and false detection.
[0085] (2) In the multi-modal fusion module of the present invention, a method combining modal weighting and cross-modal superposition is adopted. By dynamically weighting and superposing the feature information of optical images, multi-spectral images, and synthetic aperture radar images, a unified fusion feature map is generated. Combining gradient smoothing and non-linear regularization methods, the consistency and complementarity of different modal features in the high-dimensional feature space are significantly improved. By introducing a spatial attention mechanism to strengthen the key regions of the fusion feature map, the model can better capture the subtle changes and heterogeneous features in multi-modal images.
[0086] (3) The present invention enables the model to capture both the global change trend and local detail changes simultaneously at multiple scales through the dynamic diffusion method of Gaussian splash optimization of the feature map. In the change detection module, the optimized weight feature map combines multi-scale and multi-modal feature information. In cases where the change region boundary is blurred or the change type is complex, the Gaussian distribution is used to strengthen the local features, and the residual connection structure is combined to further improve the feature transfer efficiency. Compared with traditional change detection methods, the detection accuracy of the present invention in fine-grained change regions is improved by more than 10%, and it is particularly prominent in the expression of change boundaries and local features. Brief Description of the Drawings
[0087] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0088] Figure 1 is a flowchart of a remote sensing image change detection method based on a deep adaptive attention mechanism proposed by the present invention. Detailed Embodiments
[0089] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0090] Refer to Figure 1 , a remote sensing image change detection method based on a deep adaptive attention mechanism, includes the following steps:
[0091] S1. Obtain multi-temporal remote sensing image data, where the remote sensing image data includes optical images, multi-spectral images, and synthetic aperture radar images. Preprocess the remote sensing image data through radiometric correction, geometric correction, and image registration to obtain preprocessed remote sensing image data that is spatially aligned and spectrally consistent.
[0092] S2. Input the preprocessed remote sensing image data into the deep adaptive attention mechanism network model. The deep adaptive attention mechanism network model includes a feature extraction module, an attention allocation module, and a multi-modal fusion module. Use the feature extraction module to perform multi-scale feature extraction on the preprocessed remote sensing image data to generate a preliminary multi-scale feature map.
[0093] S3. Perform Gaussian splash processing on the preliminary multi-scale feature map. Based on the change center point and neighborhood distribution, diffuse the features through Gaussian distribution weights to generate a final Gaussian splash optimized feature map.
[0094] S4. Use the attention allocation module to perform dynamic weight allocation on the Gaussian splash optimized feature map, adjust the weight distribution according to the importance of the feature regions, and generate an optimized weighted feature map. The optimized weighted feature map retains the significant features of the key change regions.
[0095] S5. Input the optimized weighted feature map into the multi-modal fusion module, combine the feature information of the optical image, multi-spectral image, and synthetic aperture radar image to generate a fusion feature map. The fusion feature map represents the change information in the multi-source remote sensing image data in the form of a high-dimensional feature space.
[0096] S6. Input the fusion feature map into the change detection module to perform classification processing on the fusion feature map, identify the change regions, and generate a change detection result, including the spatial distribution of the change regions and the annotation information of the change types.
[0097] S7. Post-process the change detection result, use morphological operations to optimize the boundary continuity of the change regions and the integrity of the detection result, and finally generate an annotated remote sensing image of the change regions.
[0098] In this embodiment, S1 includes:
[0099] S11. Obtain multi-temporal remote sensing image data, where the remote sensing image data includes optical image I optical , multi-spectral image I optical , and synthetic aperture radar image I SAR ;
[0100] S12. Perform radiometric correction on the remote sensing image data.
[0101] S13. Geometrically correct the radiometrically corrected remote sensing image data by constructing a mapping function between the image and geographic coordinates. The mapping function is described using a polynomial transformation model:
[0102]
[0103] where T(x, y) represents the projection point of the image coordinate point in the geographic coordinate system, x and y represent the image coordinates respectively, and a ij is the model coefficient of geometric correction, and n is the order of the polynomial;
[0104] S14. Perform image registration on the geometrically corrected remote sensing image data. Adopt the method of feature point matching. By selecting the reference image and the image to be registered, extract the image feature point set, and use affine transformation for spatial alignment;
[0105] S15. Combine and process the remote sensing image data that has undergone radiometric correction, geometric correction, and image registration to obtain preprocessed remote sensing image data with spatial alignment and spectral consistency:
[0106] I preprocessed ={I optical,pre ,I multispectral,pre ,I SAR,pre};
[0107] where I optical,pre ,I multispectral,pre and I SAR,pre represent the preprocessed optical image, multispectral image, and synthetic aperture radar image respectively.
[0108] In this embodiment, S2 includes:
[0109] S21. Input the preprocessed remote sensing image data I preprocessed into the deep adaptive attention mechanism network model. The deep adaptive attention mechanism network model includes a feature extraction module, an attention allocation module, and a multimodal fusion module;
[0110] S22. Perform embedding mapping on different types of preprocessed remote sensing image data respectively in the feature extraction module to obtain the initial multi-scale feature representation:
[0111]
[0112] where I m,pre represents the input data of the m-th type of image, m is the image type index, including optical, multispectral, and synthetic aperture radar, l represents the multi-scale level, is the feature mapping function of the l-th layer;
[0113] S23. Input the initial multi-scale feature representation into the attention allocation module, and use the multi-head self-attention mechanism to perform deep recalibration on the features of each type of image:
[0114]
[0115] Among them, respectively represent the query, key, and value vectors for the m-th type of image feature in the l-th layer, and are the corresponding trainable weight parameters;
[0116] S24. Perform weighted operations on the query, key, and value vectors in the attention allocation module to generate a deep adaptive attention feature representation for the multi-scale features of each type of image:
[0117]
[0118] Among them, represents the attention feature that fuses local and global context information, and Attn(·) is the multi-head self-attention function, which strengthens the key regions of multi-modal data;
[0119] S25. Use the multi-modal fusion module to perform cross-modal aggregation on the deep adaptive attention features obtained from different types of images:
[0120]
[0121] to generate a preliminary multi-scale feature map:
[0122]
[0123] Among them, is a learnable weighting coefficient used to coordinate the importance of features of different types of images;
[0124] S26. Output a preliminary multi-scale feature map containing multiple scale levels:
[0125]
[0126] Among them, L is the total number of multi-scale levels.
[0127] In this embodiment, S3 includes:
[0128] S31. Input the preliminary multi-scale feature map F combined into the Gaussian splash processing module to determine the set P center ={p1, p2,..., p n} of the initial feature distribution center points of the changed regions, where p i is the coordinate of the i-th changed region center point, and n is the number of changed center points;
[0129] S32. According to the set of change center points P center Perform weight assignment on the feature image pixels around each center point. The weight is defined by a two-dimensional Gaussian distribution function as:
[0130]
[0131] where G(x, y; μ x , μ y , σ x , σ y ) is the Gaussian distribution weight, representing the weight value at position (x, y), (μ x , μ y ) represents the coordinates of the change center point, and σ x and σ y respectively represent the horizontal and vertical standard deviations of Gaussian diffusion;
[0132] S33. For the preliminary feature map at each scale l According to the Gaussian distribution weight G(x, y; μ x , μ y , σ x , σ y ) perform weighted calculation on the eigenvalues to generate a Gaussian splash optimized feature map
[0133]
[0134] where is the optimized eigenvalue at position (x, y), enhancing the feature expression in the change region and suppressing the redundant information in the background region;
[0135] S34. At each scale l, combine the weight superposition feature distributions of all change center points to calculate the final Gaussian splash optimized feature map:
[0136]
[0137] where represents the optimized feature component generated by the i-th change center point;
[0138] S35. Combine the Gaussian splash optimized feature maps at each scale l and output the set of final Gaussian splash optimized feature maps:
[0139]
[0140] In this embodiment, S4 includes:
[0141] S41. Input the set of feature maps optimized by Gaussian splash into the attention allocation module, initialize the query, key, and value vectors, and define the query vector Q(l) 、The key vector K (l) and the value vector K (l) are respectively:
[0142]
[0143] wherein, W Q(l) 、W K(l) and W V(l) are the trainable weight matrices of the l-th layer, used to generate query, key and value vectors;
[0144] S42. Calculate the attention weight matrix A (l) for the query vector Q (l) and the key vector K (l) ;
[0145] S43. Perform weighted summation on the value vector V (l) according to the attention weight matrix A (l) to generate the optimized attention feature map
[0146]
[0147] wherein, represents the optimized feature map that retains the significant features of the key change region;
[0148] S44. Introduce a residual connection to the optimized attention feature map to enhance feature transmission:
[0149]
[0150] wherein, is the final feature map that fuses the attention mechanism and Gaussian splash optimization information;
[0151] S45. Aggregate the final feature maps of all scale levels to generate an optimized set of weight feature maps:
[0152]
[0153] In this embodiment, S5 includes:
[0154] S51. Input the optimized set of weight feature maps into the multi-modal fusion module to perform modal weighting on the features of optical images, multi-spectral images and synthetic aperture radar images respectively:
[0155]
[0156] wherein, represents the fusion feature of the m-th modality at the position (x, y), is a learnable modal weighting coefficient used to adjust the contribution of each modality to feature fusion, is a gradient smoothing factor used to suppress noise in features, is the sum of the squares of the feature gradients, representing the change intensity near the position (x, y);
[0157] S53. Perform cross-modal superposition on different modal features to generate a fused feature map:
[0158]
[0159] where, represents the value of the fused feature map at the position (x, y), w m is the superposition weight of the modal feature, and ReLU(·) is an activation function used to strengthen the expression of positive features, is a non-linear regularization term;
[0160] S54. Introduce spatial attention weights to dynamically adjust the local weights of the fused feature map and calculate the attention weight matrix:
[0161]
[0162] where, S (l) (x, y) is the spatial attention weight of the fused feature map at the position (x, y), represents the adaptive scaling factor of the fused feature;
[0163] S55. Optimize the fused feature map based on the spatial attention weights to generate the final multi-modal fused feature map:
[0164]
[0165] where, is the final fused feature at the position (x, y), is the activation strengthening term.
[0166] Example 1:
[0167] In this example, an application scenario of a remote sensing image change detection technology based on a deep adaptive attention mechanism and a Gaussian splash method in a complex urban environment is described.
[0168] From July to October 2024, in order to evaluate the impact of urban expansion and infrastructure transformation on the environment, a rapidly developing coastal city in eastern China was selected as the test scene. The area contains dense urban buildings, complex road networks and a large number of green areas. At the same time, it is affected by high humidity and marine climate. There is a lot of cloud interference and spectral noise in the remote sensing images. Change detection is performed using two sets of multi-phase remote sensing image data (July and October) to identify the specific location and area of new buildings, road expansions and vegetation changes in the region.
[0169] The data comes from high-resolution optical images, multispectral images and synthetic aperture radar images, with resolutions of 0.5 meters, 5 meters and 10 meters respectively. Optical images provide detailed information, multispectral images are used for vegetation change detection, and SAR images overcome cloud and shadow interference. Each set of images covers an area of about 100 square kilometers and contains more than 2,000 ground objects.
[0170] In the data acquisition and processing stage, the remote sensing image data of July and October were preprocessed, including radiation correction, geometric correction and image registration, to ensure the spectral consistency and spatial alignment of the images. In the radiation correction stage, the areas blocked by clouds in the optical images were removed, the pixel values affected by water reflection in the multispectral images were restored, and the SAR images were denoised. In the process of geometric correction and registration, all images were aligned to a unified geographic coordinate system through ground control points and affine transformation methods.
[0171] After preprocessing, the image is input into the deep adaptive attention mechanism network model. First, the multi-scale features of optical, multispectral and SAR images are extracted in the feature extraction module, and the initial multi-scale feature maps of different modalities are generated respectively. The features of the changed areas are dynamically enhanced through Gaussian splash optimization, eliminating redundant information and background noise in the image, and eliminating the possibility of false detection caused by shadows in vegetation-covered areas. In the attention allocation module, the model automatically identifies changes in building edges, road networks and densely vegetated areas, and dynamically adjusts the feature weights according to the importance of the changed areas.
[0172] Finally, the multimodal fusion module is used to perform weighted superposition on the feature information of the three modalities to generate a fusion feature map. By combining the spectral characteristics of vegetation in multispectral images, the terrain information in SAR images, and the texture details in optical images, the model effectively identifies areas of urban expansion and environmental change.
[0173] To verify the effectiveness of the method, a comparative experiment was conducted between the method of the present invention and the traditional change detection method (change detection method based on pixel difference). In the experiment, 1,000 groups of image blocks were used as training samples, including 50 newly added buildings, 30 sections of expanded roads, and 100 vegetation change areas. The total change area was approximately 15 square kilometers. The test samples were the remaining 1,000 groups of image blocks, covering change targets with similar areas. The results of the comparative experiment are as follows:
[0174] Table 1 Comparative experiment results of the method of the present invention and the traditional change detection method
[0175]
[0176] As can be seen from Table 1, the detection accuracy of the method of the present invention is significantly higher than that of the traditional method, and the false detection rate and missed detection rate are both significantly reduced. In addition, since the model reduces redundant calculations through Gaussian splash optimization and attention allocation module, the detection time of the method of the present invention is reduced by about 40% compared with the traditional method.
[0177] Further analyze the detection results of the change area:
[0178] 1. In the detection of newly added buildings, the method of the present invention effectively identified newly added buildings with an area less than 50 square meters, and the missed detection rate of the traditional method for detecting similar targets was as high as 20%.
[0179] 2. In the detection of road expansion, the method of the present invention can accurately distinguish small road expansion areas with a width less than 2 meters. Due to the lack of the ability to model fine-grained changes in the traditional method, the false detection rate exceeds 15%.
[0180] 3. In the detection of vegetation change areas, the method of the present invention combines the characteristics of multi-spectral and SAR images to accurately distinguish vegetation reduction from seasonal spectral changes, while the traditional method is difficult to identify the actual vegetation loss areas in seasonal changes.
[0181] It can be seen from this embodiment that the method of the present invention has significant accuracy advantages and efficiency improvements in the change detection of remote sensing images in complex scenarios. The dynamic attention mechanism and multi-modal fusion ability of the model solve the problems of poor adaptability to complex backgrounds, insufficient multi-source data fusion ability, and low detection accuracy for fine-grained changes in the traditional method, providing an efficient and robust solution for the change detection of remote sensing images.
[0182] The present invention adopts a deep adaptive attention mechanism. By introducing a dynamic weight adjustment strategy in the feature extraction and allocation process, the model can focus on key change regions and suppress the interference of redundant information. The dynamic attention allocation mechanism can automatically adjust the attention weights according to the image content, strengthening the accurate recognition ability of change regions. In high vegetation coverage, complex building shadow, and high-noise scenarios, through Gaussian splash optimization of the preliminary multi-scale feature maps, the feature expression of local change regions is enhanced, solving the problems of missed detection and false detection.
[0183] The present invention adopts a method combining modal weighting and cross-modal superposition in the multi-modal fusion module. By dynamically weighting and superposing the feature information of optical images, multi-spectral images, and synthetic aperture radar images, a unified fusion feature map is generated. Combining gradient smoothing and non-linear regularization methods, the consistency and complementarity of different modal features in the high-dimensional feature space are significantly improved. By introducing a spatial attention mechanism to strengthen the key regions of the fusion feature map, the model can better capture the subtle changes and heterogeneous features in multi-modal images.
[0184] The present invention enables the model to simultaneously capture the global change trend and local detail changes at multiple scales through the dynamic diffusion method of Gaussian splash optimization of the feature map. In the change detection module, the optimized weighted feature map combines multi-scale and multi-modal feature information. When the change region boundary is blurred or the change type is complex, the Gaussian distribution is used to strengthen the local features, and the residual connection structure is combined to further improve the feature transfer efficiency. Compared with traditional change detection methods, the detection accuracy of the present invention in fine-grained change regions is improved by more than 10%, and it is particularly prominent in the change boundary and local feature expression.
[0185] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A remote sensing image change detection method based on a depth adaptive attention mechanism, characterized in that It includes the following steps: S1. Obtain multi-temporal remote sensing image data, which includes optical images, multi-spectral images, and synthetic aperture radar images. Preprocess the remote sensing image data through radiometric correction, geometric correction, and image registration to obtain preprocessed remote sensing image data with spatial alignment and spectral consistency; S2. Input the preprocessed remote sensing image data into a deep adaptive attention mechanism network model. The deep adaptive attention mechanism network model includes a feature extraction module, an attention allocation module, and a multi-modal fusion module. Use the feature extraction module to perform multi-scale feature extraction on the preprocessed remote sensing image data to generate preliminary multi-scale feature maps; S3. Perform Gaussian splash processing on the preliminary multi-scale feature maps. Based on the change center point and neighborhood distribution, diffuse the features through Gaussian distribution weights to generate final Gaussian splash optimized feature maps; S4. Use the attention allocation module to perform dynamic weight allocation on the Gaussian splash optimized feature maps, adjust the weight distribution according to the importance of the feature regions, and generate optimized weight feature maps. The optimized weight feature maps retain the significant features of the key change regions; S5. Input the optimized weight feature maps into the multi-modal fusion module, combine the feature information of optical images, multi-spectral images, and synthetic aperture radar images to generate fusion feature maps. The fusion feature maps represent the change information in the multi-source remote sensing image data in the form of a high-dimensional feature space; S6. Input the fusion feature maps into a change detection module to perform classification processing on the fusion feature maps, identify the change regions, and generate change detection results, including the spatial distribution of the change regions and the annotation information of the change types; S7. Post-process the change detection results, use morphological operations to optimize the boundary continuity of the change regions and the integrity of the detection results, and finally generate annotated remote sensing images of the change regions.
2. The remote sensing image change detection method based on a depth adaptive attention mechanism according to claim 1, wherein The S1 includes: S11. Obtain multi-temporal remote sensing image data, where the remote sensing image data includes optical image I optical , multispectral image I optical and synthetic aperture radar image I SAR ; S12. Perform radiometric correction on the relevant remote sensing image data; S13. Geometrically correct the radiometrically corrected remote sensing image data by constructing a mapping function between the image and the geographic coordinates. The mapping function is described using a polynomial transformation model: Among them, T(x, y) represents the projection point of the image coordinate point in the geographic coordinate system, x and y respectively represent the image coordinates, and a ij is the model coefficient of geometric correction, and n is the order of the polynomial; S14. Perform image registration on the geometrically corrected remote sensing image data. Adopt the method of feature point matching. By selecting a reference image and an image to be registered, extract the image feature point sets, and use affine transformation for spatial alignment; S15. Combine and process the remote sensing image data that has undergone radiometric correction, geometric correction, and image registration to obtain preprocessed remote sensing image data with spatial alignment and spectral consistency: I preprocessed ={I optical,pre ,I multispectral,pre ,I SAR,pre}; Among them, I optical,pre , I multispectral,pre and I SAR,pre respectively represent the preprocessed optical image, multispectral image, and synthetic aperture radar image.
3. A remote sensing image change detection method based on a depth adaptive attention mechanism according to claim 1, characterized in that The S2 includes: S21. Input the preprocessed remote sensing image data I preprocessed into the deep adaptive attention mechanism network model, which includes a feature extraction module, an attention allocation module, and a multimodal fusion module; S22. Perform embedding mapping on different types of preprocessed remote sensing image data respectively in the feature extraction module to obtain initial multi-scale feature representations: Among them, I m,pre represents the input data of the m-th type of image, where m is the image type index, including optical, multispectral, and synthetic aperture radar, and l represents the multi-scale level. is the feature mapping function of the l-th layer; S23. Input the initial multi-scale feature representations into the attention allocation module, and use the multi-head self-attention mechanism to perform deep recalibration on the features of each image type: wherein, respectively represent the query, key, and value vectors for the m-th type of image feature in the l-th layer, which are respectively the corresponding trainable weight parameters; S24. Perform weighted operations on the query, key, and value vectors in the attention allocation module to generate deep adaptive attention feature representations for the multi-scale features of each image type: Among them, represents the attention feature that fuses local and global context information. Attn(·) is the multi-head self-attention function, which strengthens the key regions of the multimodal data; S25. Use the multi-modal fusion module for the deep adaptive attention features obtained from different types of images: Perform cross-modal aggregation to generate a preliminary multi-scale feature map: Among them, is a learnable weighting coefficient used to coordinate the importance of features of different image types; S26. Output a preliminary multi-scale feature map containing multiple scale levels: where L is the total number of multi-scale levels.
4. A remote sensing image change detection method based on a depth adaptive attention mechanism according to claim 1, characterized in that The S3 includes: S31. Input the preliminary multi-scale feature map F combined into the Gaussian splash processing module to determine the set P center ={p1, p2, …, p n} of the initial feature distribution center points of the changed regions, where p i is the coordinate of the center point of the i-th changed region, and n is the number of changed center points; S32. According to the set P of change center points center Perform weight assignment on the feature image pixels around each center point. The weight is defined by a two-dimensional Gaussian distribution function as follows: Among them, G(x, y; μ x , μ y , σ x , σ y ) is the Gaussian distribution weight, representing the weight value at the position (x, y), (μ x , μ y ) represents the coordinates of the change center point, and σ x and σ y respectively represent the horizontal and vertical standard deviations of Gaussian diffusion; S33. Preliminary feature maps for each scale l According to the Gaussian distribution weights G(x, y; μ x , μ y , σ x , σ y ), perform weighted calculation on the eigenvalues to generate Gaussian splash optimized feature maps Among them, is the optimized eigenvalue at the position (x, y), enhancing the feature expression in the changing area and suppressing the redundant information in the background area; S34. At each scale l, calculate the final Gaussian splash optimization feature map by combining the weighted superposition feature distributions of all change center points: Among them, represents the optimized feature component generated by the i-th change center point; S35. Merge the Gaussian splash optimization feature maps of each scale l and output a set of final Gaussian splash optimization feature maps:
5. A remote sensing image change detection method based on a depth adaptive attention mechanism according to claim 1, characterized in that, The S4 includes: S41. Input the set of feature maps optimized by Gaussian splashing into the attention allocation module, initialize the query, key, and value vectors, and define the query vector Q (l) and the key vector K (l) and the value vector V (l) are respectively as follows: Among them, W Q(l) , W K(l) and W V(l) are trainable weight matrices of the l-th layer, which are used to generate query, key, and value vectors; S42. Calculate the attention weight matrix A for the query vector Q (l) and the key vector K (l) (l) ; S43. According to the attention weight matrix A (l) weight-sum the value vector V (l) to generate an optimized attention feature map Among them, represents an optimized feature map that retains the significant features of the key change region; S44. For the optimized attention feature map Introduce residual connections to enhance feature transmission: Among them, is the final feature map that integrates the attention mechanism and Gaussian splash optimization information; S45. Aggregate the final feature maps of all scale levels to generate an optimized set of weighted feature maps:
6. A remote sensing image change detection method based on a depth adaptive attention mechanism according to claim 1, characterized in that, The S5 includes: S51. Input the optimized set of weighted feature maps into a multi-modal fusion module to perform modal weighting on the features of optical images, multi-spectral images, and synthetic aperture radar images respectively: Among them, represents the fused feature of the m-th modality at the position (x, y), is a learnable modality weighting coefficient used to adjust the contribution of each modality to feature fusion, is a gradient smoothing factor used to suppress the noise in the features, is the sum of the squares of the feature gradients, representing the change intensity near the position (x, y); S53. Perform cross-modal superposition on different modal features to generate a fused feature map: Among them, represents the value of the fused feature map at position (x, y), and w m is the superimposed weight of the modal features, ReLU(·) is the activation function, which is used to strengthen the expression of positive features, and η·log(1 + is the non-linear regularization term; S54. Introduce spatial attention weights to dynamically adjust the local weights of the fused feature map and calculate the attention weight matrix: Among them, S (l) (x, y) is the spatial attention weight of the fused feature map at position (x, y), represents the adaptive scaling factor of the fused features; S55. Optimize the fused feature map based on the spatial attention weights to generate the final multi-modal fusion feature map: Among them, is the final fused feature at the position (x, y), is the activation enhancement term.