Remote sensing image change detection method and device, electronic equipment, storage medium and computer product
By using an adjusted twin neural network in remote sensing image change detection, combining pre-trained visual basic model and multi-stage semantic coding, the fusion of multi-scale semantic coding is used to guide the fusion of multi-scale semantic coding, the problem of artifacts introduced by the encoder-decoder structure is solved, and the accuracy of detection is improved.
Patent Information
- Application Number
- CN202510568170.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the existing remote sensing image change detection technology, the encoder-decoder structure is prone to introduce artifacts when processing high-resolution remote sensing images, resulting in false changes or noise in the detection results, reducing the accuracy of the detection.
A remote sensing image change detection method based on an adjusted twin neural network is proposed. By combining at least two encoder branches and a decoder, combining pre-trained visual basic model and multi-stage semantic coding, the fusion of multi-scale semantic coding is used to guide the fusion of multi-scale semantic coding to reduce artifact problems.
It significantly reduces artifact problems during the change detection process, improves the accuracy of remote sensing image change detection, and enhances the model's detection ability of small targets and edge changes.
Smart Images

Figure CN120088673A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to a remote sensing image change detection method, device, electronic device, storage medium, and computer product. Background Art
[0002] Remote sensing image change detection is an important application in remote sensing technology, aiming to identify and quantify changes in land cover or ground objects by analyzing multi-temporal remote sensing images. It can help monitor various phenomena such as land use changes, urban expansion, deforestation, natural disaster impacts, and glacier ablation, providing a scientific basis for environmental monitoring, resource management, urban planning, and disaster response.
[0003] In the process of remote sensing change detection, it is usually necessary to process multi-temporal optical images, synthetic aperture radar images, or other sensor data. Since images at different time points may be affected by factors such as sensor characteristics, imaging conditions, and illumination, change detection needs to overcome these differences and accurately extract change information.
[0004] In recent years, with the rapid development of deep learning technology, change detection methods based on deep neural networks have gradually become the mainstream. Deep learning models can automatically extract complex features in images and optimize detection performance by learning a large amount of labeled data.
[0005] In practical applications, remote sensing image change detection technology can help monitor land use changes, deforestation, urban expansion, natural disaster impacts, etc., providing important support for environmental management and decision-making. In existing remote sensing image change detection technologies, the wide application of the encoder-decoder structure has brought significant performance improvements, but at the same time, some problems have also emerged, especially the artifact problem. When processing high-resolution remote sensing images, this structure is prone to introducing artifacts due to its inherent upsampling and downsampling processes, resulting in false changes or noise in the detection results, interfering with the identification of real changes, and reducing the accuracy of remote sensing image change detection. Summary of the Invention
[0006] This application aims to at least solve one of the technical problems existing in the related art. For this purpose, this application proposes a remote sensing image change detection method, device, electronic device, storage medium, and computer product to solve the problem of easily introducing artifacts in existing remote sensing image change detection technologies and achieve improved accuracy of remote sensing image change detection.
[0007] The remote sensing image change detection method according to the first aspect embodiment of this application includes: Obtain two-temporal remote sensing images to be detected; Input two - phase remote sensing images to be detected into the change detection model, and obtain the change detection result output by the change detection model; wherein, the change detection model is constructed based on the adjusted Siamese neural network; the adjusted Siamese neural network includes at least two encoder branches and a decoder; each encoder branch is respectively used to extract the multi - scale semantic encoding of the input image by combining the first branch and the second branch therein; the decoder is used to fuse and encode the multi - scale semantic encodings respectively extracted by each encoder branch through high - and low - frequency information guidance to obtain the result of change detection; the first branch is a pre - trained visual foundation model; the second branch is used to perform multi - stage semantic encoding on the input image, and the semantic encoding of each stage is respectively guided by high - and low - frequency information.
[0008] According to an embodiment of the present application, the second branch is specifically used for: Extract features from the input image to obtain the initial component; Perform frequency - domain decoupling on the initial component to obtain the high - frequency component and the low - frequency component; Based on the initial component, the high - frequency component and the low - frequency component, and in combination with the feature components extracted by the first branch, perform multi - stage semantic encoding.
[0009] According to an embodiment of the present application, the performing multi - stage semantic encoding based on the initial component, the high - frequency component and the low - frequency component, and in combination with the feature components extracted by the first branch includes: Perform cross - attention calculation based on the high - frequency component and the initial component to obtain the first component; Perform information gain on the first component based on the feature components extracted by the first branch to obtain the second component; Perform cross - attention calculation based on the high - frequency component and the second component to obtain the third component; Perform information gain on the third component based on the feature components extracted by the first branch to obtain the fourth component; Perform cross - attention calculation based on the low - frequency component and the fourth component to obtain the fifth component; Perform information gain on the fifth component based on the feature components extracted by the first branch to obtain the sixth component; Perform cross - attention calculation based on the low - frequency component and the sixth component to obtain the seventh component.
[0010] According to an embodiment of the present application, the decoder is specifically used for: Perform feature fusion on the first component and the third component respectively extracted by each encoder branch to obtain the first fusion feature; Feature fusion is performed on the fifth component and the seventh component respectively extracted from each encoder branch to obtain a second fused feature; Feature fusion is performed on the fifth component and the seventh component respectively extracted from each encoder branch to obtain a third fused feature; Feature fusion is performed on the first component, the third component, the fifth component and the seventh component respectively extracted from each encoder branch to obtain a fourth fused feature; Feature fusion is performed on the first fused feature and the fourth fused feature to obtain a fifth fused feature; Feature fusion is performed on the second fused feature and the third fused feature to obtain a sixth fused feature; Change detection is performed on the fifth fused feature and the sixth fused feature to obtain the result of change detection.
[0011] According to an embodiment of the present application, the feature fusion of the first component and the third component respectively extracted from each encoder branch to obtain a first fused feature includes: Differential fusion is performed on the first components respectively extracted from each encoder branch to obtain a first differential fused feature; Differential fusion is performed on the third components respectively extracted from each encoder branch to obtain a second differential fused feature; Scale alignment is performed on the first differential fused feature and the second differential fused feature to obtain a first fused feature.
[0012] According to an embodiment of the present application, the scale alignment of the first differential fused feature and the second differential fused feature to obtain a first fused feature includes: Upsampling processing is performed on the second differential fused feature to obtain a third differential fused feature; Feature splicing is performed on the third differential fused feature and the first differential fused feature to obtain a first fused feature.
[0013] According to the remote sensing image change detection device of the second aspect embodiment of the present application, it includes: An acquisition module, configured to acquire two-phase remote sensing images to be detected; A detection module is configured to input two-phase remote sensing images to be detected into a change detection model, and obtain a change detection result output by the change detection model. Among them, the change detection model is constructed based on an adjusted Siamese neural network. The adjusted Siamese neural network includes at least two encoder branches and a decoder. Each encoder branch is respectively configured to extract multi-scale semantic encodings of the input image by combining a first branch and a second branch therein. The decoder is configured to fuse and encode the multi-scale semantic encodings respectively extracted by each encoder branch through high-frequency and low-frequency information guidance to obtain a change detection result. The first branch is a pre-trained visual foundation model. The second branch is configured to perform multi-stage semantic encoding on the input image, and the semantic encoding of each stage is respectively guided by high-frequency and low-frequency information.
[0014] An electronic device according to an embodiment of the third aspect of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the remote sensing image change detection method as described in any one of the above is implemented.
[0015] A storage medium according to an embodiment of the fourth aspect of the present application is a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the remote sensing image change detection method as described in any one of the above is implemented.
[0016] A computer program product according to an embodiment of the fifth aspect of the present application includes a computer program. When the computer program is executed by a processor, the remote sensing image change detection method as described in any one of the above is implemented.
[0017] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects: A change detection model is constructed through an adjusted Siamese neural network. Since the adjusted Siamese neural network includes at least two encoder branches and one decoder. Each encoder branch is used to extract the multi-scale semantic encoding of the input image by combining a first branch which is a pre-trained visual base model therein, and a second branch for multi-stage semantic encoding of the input image; while the decoder is used to fuse and encode the multi-scale semantic encodings separately extracted by each encoder branch through high and low frequency information guidance to obtain the result of change detection. Therefore, after obtaining the two-phase remote sensing images to be detected, inputting the two-phase remote sensing images to be detected into the change detection model, the change detection result output by the change detection model can be obtained. Since the semantic encoding of each stage is guided by high and low frequency information respectively, and the multi-scale semantic encodings are also fused through high and low frequency information guidance, making full use of the fusion and indication information of each stage, weakening the problems of high frequency information failing in the deep layer and low frequency information being blurred at the edge, thereby significantly reducing the artifact problem in the change detection process and achieving the improvement of the accuracy of remote sensing image change detection.
[0018] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flowchart of a remote sensing image change detection method provided by an embodiment of the present application.
[0021] Figure 2 It is a schematic structural diagram of a decoder in a remote sensing image change detection method provided by an embodiment of the present application.
[0022] Figure 3 It is a schematic structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following further describes the embodiments of the present application in detail in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.
[0024] In the description of the embodiments of the present application, it should be noted that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the embodiments of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0025] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "connected" and "connected to" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.
[0026] In the embodiments of the present application, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature can be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0027] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0028] It should be noted that the artifact problem is also related to the complexity of data preprocessing. The input of multi-source heterogeneous remote sensing data increases the difficulty of preprocessing, such as high-precision correction and registration of multi-resolution images. The improvement of image spatial resolution and data quality also makes the pseudo-changes caused by imaging environmental factors such as illumination, terrain, and shadow more significant. These problems not only affect the accuracy of change detection, but also pose challenges to the generalization ability of the model and the actual application effect.
[0029] Specifically, the artifact regions often appear near the edge transition regions between the specific change regions and the unchanged regions. The low-confidence edges caused by the upsampling strategy are the fundamental reasons for the generation of artifacts. In the pixel-level change detection task based on deep learning, the artifact interference problem is an important bottleneck restricting the model performance. The specific manifestations are as follows: When upsampling and restoring the deep features extracted by the encoder, in the spatially discontinuous regions such as object boundaries and texture mutations (i.e., the transition zones between the regions where geographical entities change and the adjacent unchanged regions), strip-shaped or patchy abnormal response patterns often occur. This phenomenon stems from the modeling defects of conventional upsampling strategies (such as bilinear interpolation and transposed convolution) for high-frequency information - during the reconstruction of the spatial dimension of the feature map, there is a lack of special processing mechanisms for edge-sensitive regions, resulting in systematic biases in the confidence evaluation of the model for the transition regions. Such low-confidence edges make it difficult to accurately coordinate the semantic consistency of different-level features during the decoding stage, and finally two typical artifacts appear in the prediction map: one is the serrated structural distortion, and the other is the diffuse salt-and-pepper noise. Experiments show that when the shape of the target object shows a complex and irregular distribution (such as the junction area between urban building groups and natural landforms), the artifact effect will be significantly exacerbated with the increase of spatial heterogeneity.
[0030] Based on this, the present application proposes a remote sensing image change detection method, device, electronic device, storage medium, and computer product.
[0031] Figure 1 It is a schematic flowchart of the remote sensing image change detection method provided by the embodiment of the present application. As Figure 1 shown, the remote sensing image change detection method includes: Step 110, obtaining two-phase remote sensing images to be detected.
[0032] Step 120: Input two temporal remote sensing images to be detected into the change detection model to obtain the change detection result output by the change detection model. Among them, the change detection model is constructed based on the adjusted Siamese neural network. The adjusted Siamese neural network includes at least two encoder branches and a decoder. Each encoder branch is respectively used to extract the multi-scale semantic encoding of the input image by combining the first branch and the second branch therein. The decoder is used to fuse and encode the multi-scale semantic encodings respectively extracted by each encoder branch through the guidance of high and low frequency information to obtain the result of change detection. The first branch is a pre-trained visual base model. The second branch is used to perform multi-stage semantic encoding on the input image, and the semantic encoding of each stage is respectively guided by high and low frequency information.
[0033] It should be noted that the execution subject of the remote sensing image change detection method provided in the embodiment of the present application can be a computer device, such as a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc. It should be noted that all the data that needs to be obtained in the present application is obtained through formal channels after being authorized by relevant users.
[0034] The remote sensing image change detection device can be set or connected in the computer device of the present application, so as to control the remote sensing image change detection device to execute the remote sensing image change detection method of the present application.
[0035] Specifically, the present application can obtain remote sensing images such as multi-temporal optical images, synthetic aperture radar images, or other sensor data.
[0036] Furthermore, one or more preprocessing operations such as registration, denoising, and normalization can be performed on the multi-temporal remote sensing images to ensure the consistency of the images in space and spectrum and provide a basis for subsequent analysis.
[0037] It should be noted that the present application can adjust the traditional Siamese neural network structure so that the adjusted Siamese neural network includes at least two encoder branches and a decoder. In a specific embodiment, the adjusted Siamese neural network can include two encoder branches and a decoder.
[0038] Among them, each encoder branch is respectively used to extract the multi-scale semantic encoding of the input image by combining the first branch and the second branch therein.
[0039] The first branch is a pre-trained Visual Foundation Model (VFM), which is specifically pre-trained with a natural image dataset and undertakes the core task of extracting general features. This model is pre-trained based on a large-scale natural image dataset and can capture rich visual semantic information, thereby providing high-quality initial feature representations for subsequent tasks.
[0040] Specifically, the first branch can be Clip_Vit_Base, which is a multi-modal pre-trained model based on the Contrastive Language-Image Pretraining (CLIP) framework. Its design combines the architectural advantages of the Vision Transformer (ViT), effectively modeling global dependencies in images through the self-attention mechanism. At the same time, by leveraging the knowledge learned during the pre-training process, it significantly improves the model's generalization ability in various downstream tasks. The natural image dataset used in the pre-training stage usually contains millions of annotated images, covering a wide range of visual scenes and object categories, which enables Clip_Vit_Base to learn general visual features and be applicable to various computer vision tasks such as image classification, object detection, and semantic segmentation. By using Clip_Vit_Base as the basic feature extractor, the dependence on data annotation in downstream tasks can be significantly reduced, while improving the performance and robustness of the model.
[0041] Therefore, the input image can be feature-extracted through the first branch to obtain the feature components of the input image.
[0042] In this application, the second branch can be defined as an expert network, which is used to perform multi-stage semantic encoding on the input image, and the semantic encoding at each stage is guided by high-frequency and low-frequency information respectively.
[0043] Therefore, multi-scale semantic encodings extracted by the two encoder branches for the input image can be obtained, that is, each encoder branch extracts multi-scale semantic encodings.
[0044] The decoder is used to fuse and encode the multi-scale semantic encodings extracted by each encoder branch through the guidance of high-frequency and low-frequency information to obtain the result of change detection.
[0045] In this application, the decoder can use a frequency-domain information fusion decoder based on a Multilayer Perceptron (MLP) to achieve feature aggregation of multiple encoder branches and generation of the prediction map. Its core advantage lies in that through an efficient feature extraction and aggregation mechanism, it can significantly reduce the computational cost while maintaining high accuracy.
[0046] The MLP-based frequency-domain information fusion decoder effectively integrates feature maps of different scales through hierarchical feature fusion, thus achieving a balance between semantic information and spatial details. This design not only improves the model's detection ability for small targets and edge changes but also enhances its robustness to different scale changes by combining a deep supervision strategy with multi-scale prediction maps.
[0047] Furthermore, a change detection model is constructed based on the adjusted Siamese neural network and trained using the collected multi-temporal remote sensing data as training data, thereby realizing the training of the expert network therein. After training, the change detection model can detect whether there are changes in the input two-temporal remote sensing images and output the change detection results.
[0048] Thus, when remote sensing image change detection is required, the two-temporal remote sensing images to be detected can be input into the change detection model, and then the change detection model can perform change detection and obtain the change detection results output by the change detection model. Among them, before inputting into the change detection model, one or more of the above preprocessing operations can be performed on the two-temporal remote sensing images to be detected.
[0049] Furthermore, according to the change detection results, various phenomena such as monitoring land use changes, urban expansion, deforestation, natural disaster impacts, and glacier ablation can be realized according to specific scenarios, providing a scientific basis for environmental monitoring, resource management, urban planning, and disaster response.
[0050] According to the remote sensing image change detection method of the embodiments of the present application, a change detection model is constructed through the adjusted Siamese neural network. Since the adjusted Siamese neural network includes at least two encoder branches and one decoder. Each encoder branch is used to extract the multi-scale semantic encoding of the input image by combining the first branch, which is a pre-trained visual foundation model, and the second branch, which is used for multi-stage semantic encoding of the input image; while the decoder is used to fuse and encode the multi-scale semantic encodings separately extracted by each encoder branch through high-frequency and low-frequency information guidance to obtain the change detection results. Therefore, after obtaining the two-temporal remote sensing images to be detected, inputting the two-temporal remote sensing images to be detected into the change detection model can obtain the change detection results output by the change detection model. Since the semantic encoding of each stage is guided by high-frequency and low-frequency information respectively, and the multi-scale semantic encoding is also fused through high-frequency and low-frequency information guidance, making full use of the fusion and indication information of each stage, weakening the problems of high-frequency information becoming invalid in the deep layer and low-frequency information being blurred at the edge, thus significantly reducing the artifact problem in the change detection process and realizing the improvement of the accuracy of remote sensing image change detection.
[0051] Based on the above embodiments, the second branch is specifically used for: Performing feature extraction on the input image to obtain an initial component; Perform frequency-domain decoupling on the initial component to obtain a high-frequency component and a low-frequency component; Based on the initial component, the high-frequency component, and the low-frequency component, combined with the feature components extracted by the first branch, perform multi-stage semantic encoding.
[0052] Specifically, a frequency-domain decoupling module DeCoupe is set in the second branch of the present application, which can obtain high and low frequency components for an input image at any certain time phase. Among them, to obtain the high and low frequency components, the low-frequency component can be extracted through convolution and pooling operations, and then the high-frequency component can be obtained by subtracting the low-frequency component from the original component.
[0053] Specifically, the frequency-domain decoupling module can be used to first perform a convolution operation on the input image to capture local features and obtain the initial component; then, the low-frequency component is extracted from the initial component through pooling (such as average pooling or max pooling) to retain the global structure; finally, the high-frequency component is obtained by subtracting the low-frequency component from the original component, which reflects details and edge information. This decomposition method can effectively separate the global and local information in the features and is suitable for tasks such as image processing and change detection. By combining the high-frequency and low-frequency components, the module can enhance the feature expression ability and improve the model performance.
[0054] Based on this, further multi-stage semantic encoding can be performed according to the initial component, the high-frequency component, and the low-frequency component, combined with the feature components extracted by the first branch.
[0055] Specifically, in the present application, initial-stage semantic encoding can be first performed based on the initial component and the high-frequency component. Specifically, an encoding module based on the SegFormer structure can be used, introducing the DeFreqTransformer Block, which realizes guided learning of the features of the original image in high and low frequency stages. SegFormer is a semantic segmentation model based on Transformer, which aims to solve the challenges faced by traditional convolutional neural networks in semantic segmentation tasks.
[0056] It should be noted that during the forward process of the DeFreqTransformer Block, the initial component x at each stage will go through two attention stages. First, self-attention is performed, which is the standard Transformer attention mechanism. After self-attention, x will go through cross-attention guided by the frequency-domain component (i.e., the high-frequency component or the low-frequency component) x_freq at each stage, that is, using x_freq as the query, and x as the key and value to complete the cross-attention calculation.
[0057] It should be noted that the output of each stage serves as the initial component of the next stage.
[0058] It should be further noted that after each coding stage of this application, the second branch can use the bridge module in the BAN framework to communicate with the visual base model. Specifically, through a simple cross-attention method, information gain can be performed on the components of the original information (i.e., the initial components of the next stage) based on the feature components extracted by the visual base model. In terms of the data flow direction, the original information passes through both the expert network and the visual base model synchronously. Among them, the BAN framework is a Bilinear Attention Network (abbreviated as BAN) framework.
[0059] Furthermore, multi-stage semantic coding can be completed based on the initial components and high-frequency components to obtain a specified number of components in the high-frequency stage.
[0060] Further, after the components in the last high-frequency stage communicate with the visual base model, the obtained components are combined with the low-frequency components for multi-stage semantic coding to obtain a specified number of components in the low-frequency stage.
[0061] In a specific embodiment, each encoder branch can perform two-stage high-frequency semantic coding and two-stage low-frequency semantic coding on the input image, and finally obtain two components in the high-frequency stage and two components in the low-frequency stage.
[0062] The two encoder branches can obtain a total of four components in the high-frequency stage and four components in the low-frequency stage.
[0063] This application completes the guidance of the attention focus of the expert network by selecting suitable high- and low-frequency information in different downsampling stages, weakens the problems of high-frequency information failure in the deep layer and low-frequency information blurring at the edge, helps to reduce the artifact problem in the change detection process, and improves the accuracy of remote sensing image change detection.
[0064] Based on the above embodiments, multi-stage semantic coding is performed based on the initial components, high-frequency components, and low-frequency components, combined with the feature components extracted by the first branch, including: Performing cross-attention calculation based on the high-frequency components and the initial components to obtain the first component; Performing information gain on the first component based on the feature components extracted by the first branch to obtain the second component; Performing cross-attention calculation based on the high-frequency components and the second component to obtain the third component; Performing information gain on the third component based on the feature components extracted by the first branch to obtain the fourth component; Performing cross-attention calculation based on the low-frequency components and the fourth component to obtain the fifth component; Perform information gain on the fifth component based on the feature components extracted from the first branch to obtain the sixth component; Perform cross-attention calculation based on the low-frequency component and the sixth component to obtain the seventh component.
[0065] Specifically, this application can utilize the DeFreqTransformer Block to perform cross-attention calculation based on the high-frequency component and the initial component to obtain the first component. Specifically, self-attention can be first performed on the initial component x, and this operation is a standard transformer attention mechanism. The initial component x after self-attention will go through cross-attention guided by the frequency-domain component x_freq at each stage. Specifically, the initial component x after self-attention can be used as the key and value, and the high-frequency component x_freq can be used as the query for cross-attention calculation, and the result of the cross-attention is defined as the first component. The first component is the component of the first high-frequency stage.
[0066] Furthermore, utilize the bridge module to communicate with the visual base model, and then through a simple cross-attention method, perform information gain on the first component based on the feature components extracted by the visual base model to obtain the second component.
[0067] Furthermore, based on the same method as the first high-frequency stage, perform cross-attention calculation based on the high-frequency component and the second component to obtain the third component. The third component is the component of the second high-frequency stage.
[0068] Furthermore, based on the same information gain method, perform information gain on the third component based on the feature components extracted from the first branch to obtain the fourth component.
[0069] Furthermore, based on the same method as the second high-frequency stage, perform cross-attention calculation based on the low-frequency component and the fourth component to obtain the fifth component. The fifth component is the component of the first low-frequency stage.
[0070] Furthermore, based on the same information gain method, perform information gain on the fifth component based on the feature components extracted from the first branch to obtain the sixth component.
[0071] Furthermore, based on the same method as the first low-frequency stage, perform cross-attention calculation based on the low-frequency component and the sixth component to obtain the seventh component. The seventh component is the component of the second low-frequency stage.
[0072] Thus, the multi-stage semantic encoding corresponding to this encoder branch can be completed to obtain the corresponding number of components.
[0073] Furthermore, the components of each encoder branch can be input into the decoder for subsequent processing.
[0074] By selecting suitable high - and low - frequency information at different downsampling stages in this application, the guidance of the attention focus of the expert network is completed, weakening the problems of high - frequency information failure in the deep layer and low - frequency information blurring at the edge, which helps to reduce the artifact problem in the change detection process and improve the accuracy of remote - sensing image change detection.
[0075] Based on the above - mentioned embodiments, the decoder is specifically used for: Performing feature fusion on the first component and the third component respectively extracted by each encoder branch to obtain a first fused feature; Performing feature fusion on the fifth component and the seventh component respectively extracted by each encoder branch to obtain a second fused feature; Performing feature fusion on the fifth component and the seventh component respectively extracted by each encoder branch to obtain a third fused feature; Performing feature fusion on the first component, the third component, the fifth component, and the seventh component respectively extracted by each encoder branch to obtain a fourth fused feature; Performing feature fusion on the first fused feature and the fourth fused feature to obtain a fifth fused feature; Performing feature fusion on the second fused feature and the third fused feature to obtain a sixth fused feature; Performing change detection on the fifth fused feature and the sixth fused feature to obtain the result of change detection.
[0076] Figure 2 It is a schematic structural diagram of the decoder in the remote - sensing image change - detection method provided by the embodiments of this application. As Figure 2 shown, a high - frequency differential feed - forward network High_dis FFN, a low - frequency differential feed - forward network Low_dis FFN, a staged differential feed - forward network Staged FFN, a high - frequency feature fusion network High merge, a low - frequency feature fusion network Low merge, and a discriminator Discriminator can be set in the decoder of this application.
[0077] Therefore, the high - frequency differential feed - forward network High_dis FFN can be used to perform feature fusion on the first component I and the third component II respectively extracted by each encoder branch to obtain a first fused feature.
[0078] The low - frequency differential feed - forward network Low_dis FFN is used to perform feature fusion on the fifth component III and the seventh component IV respectively extracted by each encoder branch to obtain a second fused feature.
[0079] The staged differential feed - forward network Staged FFN is used to perform feature fusion on the fifth component and the seventh component respectively extracted by each encoder branch to obtain a third fused feature.
[0080] Using the Staged Feed-Forward Network (Staged FFN), feature fusion is performed based on the first, third, fifth, and seventh components separately extracted by each encoder branch to obtain the fourth fused feature.
[0081] It should be noted that the above differential feed-forward networks have similar structures. Essentially, they all add positional encoding through a convolutional layer on the basis of a multi-layer perceptron to achieve better learning effects and generalization capabilities. The High-Frequency Differential Feed-Forward Network processes the encoded components in the high-frequency stage, and the principles of the other two differential feed-forward networks are the same. In this part, the decoder obtains the inputs from two time phases for the first time and combines them to perform a differential operation. Essentially, it is from the decoder that the network truly begins to perceive changes.
[0082] Furthermore, using the High-Frequency Feature Fusion Network (High merge), feature fusion is performed based on the first fused feature and the fourth fused feature to obtain the fifth fused feature.
[0083] Specifically, through the High-Frequency Feature Fusion Network (High merge), feature concatenation (Concate) can be performed on the first fused feature and the fourth fused feature, and the concatenated feature is defined as the fifth fused feature.
[0084] Using the Low-Frequency Feature Fusion Network (Low merge), feature fusion is performed based on the second fused feature and the third fused feature to obtain the sixth fused feature.
[0085] Specifically, the Low-Frequency Feature Fusion Network (Low merge) can be used to perform feature concatenation (Concate) on the second fused feature and the third fused feature, and the concatenated feature is defined as the sixth fused feature.
[0086] Finally, using the Discriminator, change detection is performed based on the fifth fused feature and the sixth fused feature to obtain the result of change detection. Specifically, the fifth fused feature and the sixth fused feature can be input into the Discriminator to obtain the result of change detection output by the Discriminator. Among them, the result of change detection can specifically be a binary image. Different values in the image indicate whether there are changes in the corresponding pixel points.
[0087] This application redesigned the frequency-domain fusion decoder suitable for the frequency-domain information stage guidance network, making full use of the fusion and indication information at each stage. Specifically, the decoder part processes the multi-scale semantic encoded information through the High-frequency differential feed-forward network (High_dis FFN), the Low-frequency differential feed-forward network (Low_dis FFN), and the Staged feed-forward network (Staged FFN), and extracts high-frequency, low-frequency, and stage differential features respectively. These features are effectively integrated through a fusion method guided by high- and low-frequency information, making full use of the fusion and indication information at each stage, weakening the problems of high-frequency information becoming ineffective in the deep layer and low-frequency information being blurred at the edge, helping to reduce the artifact problem in the change detection process, and improving the accuracy of remote sensing image change detection.
[0088] This application also improves the model's detection ability for small targets, edge changes, and different-scale changes through multi-scale feature fusion and deep supervision strategies, achieving a balance between high accuracy and high efficiency.
[0089] Based on the above embodiments, feature fusion is performed on the first component and the third component respectively extracted from each encoder branch to obtain the first fusion feature, including: Perform differential fusion on the first components respectively extracted from each encoder branch to obtain the first differential fusion feature; Perform differential fusion on the third components respectively extracted from each encoder branch to obtain the second differential fusion feature; Align the scales of the first differential fusion feature and the second differential fusion feature to obtain the first fusion feature.
[0090] Specifically, taking the use of the High-frequency differential feed-forward network High_dis FFN to perform feature fusion on the first component and the third component respectively extracted from each encoder branch as an example, the process of performing feature fusion using the differential feed-forward network is described.
[0091] Specifically, the High-frequency differential feed-forward network High_dis FFN can be used to perform differential fusion on the first components respectively extracted from each encoder branch to obtain the first differential fusion feature. For example, perform differential fusion on the first components extracted from two encoder branches, and define the differential fusion result of the two first components as the first differential fusion feature.
[0092] More specifically, the difference operation can be performed on the two first components, and the result of the difference operation is the first differential fusion feature.
[0093] Based on the same principle, the High-frequency differential feed-forward network High_dis FFN is used to perform differential fusion on the third components respectively extracted from each encoder branch to obtain the second differential fusion feature.
[0094] Furthermore, the High_dis FFN (High-frequency Differential Feed-forward Network) can be used to align the scales of the first differential fusion feature and the second differential fusion feature, and the first fusion feature can be obtained after the scale alignment is completed.
[0095] In this application, the multi-scale semantic encoding information is processed by the High-frequency Differential Feed-forward Network (High_dis FFN), the Low-frequency Differential Feed-forward Network (Low_disFFN), and the Staged Feed-forward Network (Staged FFN) to extract high-frequency, low-frequency, and stage differential features respectively. These features are effectively integrated through a fusion method guided by high and low frequency information, making full use of the fusion and indication information at each stage, weakening the problems of high-frequency information becoming ineffective in the deep layer and low-frequency information being blurred at the edges, helping to reduce the artifact problem in the change detection process, and improving the accuracy of remote sensing image change detection.
[0096] Based on the above embodiments, aligning the scales of the first differential fusion feature and the second differential fusion feature to obtain the first fusion feature includes: Performing an upsampling process on the second differential fusion feature to obtain a third differential fusion feature; Performing feature concatenation on the third differential fusion feature and the first differential fusion feature to obtain the first fusion feature.
[0097] Specifically, after the first differential fusion feature and the second differential fusion feature are obtained, since the scales of the two differential fusion features are different, the second differential fusion feature with a smaller scale can be first upsampled so that the scale of the obtained third differential fusion feature is the same as the scale of the first differential fusion feature. Then, feature concatenation can be performed on the third differential fusion feature and the first differential fusion feature, and the first fusion feature can be obtained after the concatenation is completed.
[0098] In this application, by aligning features of different scales to fuse high-frequency and low-frequency features, the model's ability to capture details and global information is enhanced, making it show stronger robustness and adaptability in complex scenarios, which helps to improve the accuracy of remote sensing image change detection.
[0099] In this application, by guiding the neural network attention mechanism in stages, the artifact problem in the low-confidence area during the change detection process is reduced, that is, the features decoded in the high-resolution small receptive field part are guided by high-frequency information for sharp edge segmentation, and the features decoded in the low-resolution large receptive field part are guided by low-frequency information for block segmentation and qualitative analysis.
[0100] Furthermore, this application can conduct comparative experiments based on the LEVIR-CD dataset. Among them, the LEVIR-CD dataset is one of the recognized important benchmark datasets in the field of remote sensing change detection. ChangeFormer and BAN are milestone technologies in the field of remote sensing change detection. Both of these two technologies and this application adopt dual-temporal input to obtain binary output. Under fair and identical conditions, this application is compared with the above two technologies on the LEVIR-CD dataset.
[0101] The experimental results are shown in Table 1 below: Table 1 Comparison of the accuracy of DeFreq before and after improvement on the LEVIR-CD dataset
[0102] The F1 score is the harmonic mean of precision and recall, which is used to comprehensively measure the accuracy and integrity of the model; the Intersection over Union (IoU) represents the degree of overlap between the predicted region and the true region, which is used to evaluate the accuracy of segmentation or detection tasks; precision (Precision, Pre) measures the proportion of samples predicted as positive by the model that are actually positive, reflecting the reliability of the model's prediction.
[0103] This application is superior to the current advanced technology BAN in all major indicators, achieving IoU +0.77%, F1 +0.45%, and Pre +1.37% respectively, and is at a leading level. The experiment shows that the phased frequency domain attention guidance mechanism can improve the accuracy.
[0104] Next, the remote sensing image change detection device provided by this application will be described. The remote sensing image change detection device described below can be mutually referred to the remote sensing image change detection method described above.
[0105] Furthermore, this application also provides a remote sensing image change detection device.
[0106] The remote sensing image change detection device includes: An acquisition module for acquiring two-temporal remote sensing images to be detected; A detection module, configured to input two-phase remote sensing images to be detected into a change detection model, and obtain a change detection result output by the change detection model; wherein, the change detection model is constructed based on an adjusted Siamese neural network; the adjusted Siamese neural network includes at least two encoder branches and a decoder; each encoder branch is respectively configured to extract multi-scale semantic encodings of the input image by combining a first branch and a second branch therein; the decoder is configured to fuse and encode the multi-scale semantic encodings respectively extracted by each encoder branch through high-frequency and low-frequency information guidance to obtain a change detection result; the first branch is a pre-trained visual base model; the second branch is configured to perform multi-stage semantic encoding on the input image, and the semantic encoding of each stage is respectively guided by high-frequency and low-frequency information.
[0107] The remote sensing image change detection device of the present application constructs a change detection model through an adjusted Siamese neural network. Since the adjusted Siamese neural network includes at least two encoder branches and a decoder. Each encoder branch is configured to extract multi-scale semantic encodings of the input image by combining a first branch which is a pre-trained visual base model and a second branch which is configured to perform multi-stage semantic encoding on the input image; and the decoder is configured to fuse and encode the multi-scale semantic encodings respectively extracted by each encoder branch through high-frequency and low-frequency information guidance to obtain a change detection result. Therefore, after obtaining two-phase remote sensing images to be detected, input the two-phase remote sensing images to be detected into the change detection model, and the change detection result output by the change detection model can be obtained. Since the semantic encoding of each stage is respectively guided by high-frequency and low-frequency information, and the multi-scale semantic encodings are also fused through high-frequency and low-frequency information guidance, making full use of the fusion and indication information of each stage, weakening the problems of high-frequency information failing in the deep layer and low-frequency information being blurred at the edge, thereby significantly reducing the artifact problem in the change detection process and achieving improved accuracy of remote sensing image change detection.
[0108] Figure 3 An example of the physical structure diagram of an electronic device is shown as Figure 3 shown. The electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 310 may call logic instructions in the memory 330 to execute the following method: obtain two-phase remote sensing images to be detected; Input two - phase remote sensing images to be detected into the change detection model to obtain the change detection result output by the change detection model; wherein, the change detection model is constructed based on an adjusted Siamese neural network; the adjusted Siamese neural network includes at least two encoder branches and a decoder; each encoder branch is respectively used to extract multi - scale semantic encodings of the input image by combining the first branch and the second branch therein; the decoder is used to fuse and encode the multi - scale semantic encodings respectively extracted by each encoder branch through the guidance of high - and low - frequency information to obtain the change detection result; the first branch is a pre - trained visual basic model; the second branch is used to perform multi - stage semantic encoding on the input image, and the semantic encoding of each stage is respectively guided by high - and low - frequency information.
[0109] In addition, when the logical instructions in the above - mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer - readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read - only memories (ROM, Read - Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0110] In another aspect, an embodiment of this application also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above - mentioned embodiments, for example, including: obtaining two - phase remote sensing images to be detected; Input two - phase remote sensing images to be detected into the change detection model to obtain the change detection result output by the change detection model; wherein, the change detection model is constructed based on an adjusted Siamese neural network; the adjusted Siamese neural network includes at least two encoder branches and a decoder; each encoder branch is respectively used to extract multi - scale semantic encodings of the input image by combining the first branch and the second branch therein; the decoder is used to fuse and encode the multi - scale semantic encodings respectively extracted by each encoder branch through the guidance of high - and low - frequency information to obtain the change detection result; the first branch is a pre - trained visual basic model; the second branch is used to perform multi - stage semantic encoding on the input image, and the semantic encoding of each stage is respectively guided by high - and low - frequency information.
[0111] In another aspect, an embodiment of the present application further provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above embodiments, for example, including: obtaining two-phase remote sensing images to be detected; Inputting the two-phase remote sensing images to be detected into a change detection model to obtain a change detection result output by the change detection model; wherein, the change detection model is constructed based on an adjusted Siamese neural network; the adjusted Siamese neural network includes at least two encoder branches and a decoder; each encoder branch is respectively configured to extract multi-scale semantic encodings of the input image by combining a first branch and a second branch therein; the decoder is configured to fuse and encode the multi-scale semantic encodings respectively extracted by each encoder branch through high-frequency and low-frequency information guidance to obtain a change detection result; the first branch is a pre-trained visual foundation model; the second branch is configured to perform multi-stage semantic encoding on the input image, and the semantic encoding at each stage is respectively guided by high-frequency and low-frequency information.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application.
Claims
1. A remote sensing image change detection method, characterized in that: include: Acquire two-phase remote sensing images to be detected; Input the two-phase remote sensing images to be detected into the change detection model to obtain the change detection result output by the change detection model; wherein the change detection model is constructed based on the adjusted twin neural network; the adjusted twin neural network includes at least two encoder branches and one decoder; each encoder branch is used to extract the multi-scale semantic coding of the input image by combining the first branch and the second branch therein; the decoder is used to fuse the multi-scale semantic coding extracted by each encoder branch, and obtain the change detection result by encoding after fusing through the guidance of high and low frequency information; The first branch is a pre-trained visual basic model; the second branch is used to perform multi-stage semantic encoding on the input image, and the semantic encoding of each stage is guided by high-frequency and low-frequency information respectively.
2. The remote sensing image change detection method according to claim 1, characterized in that: The second branch is specifically used for: Perform feature extraction based on the input image to obtain initial components; Decoupling the initial component in the frequency domain to obtain a high-frequency component and a low-frequency component; Based on the initial component, the high-frequency component and the low-frequency component, multi-stage semantic encoding is performed in combination with the feature component extracted by the first branch.
3. The remote sensing image change detection method according to claim 2, characterized in that: The multi-stage semantic coding based on the initial component, the high-frequency component and the low-frequency component combined with the feature component extracted by the first branch includes: Performing cross-attention calculation based on the high-frequency component and the initial component to obtain a first component; Performing information gain on the first component based on the feature component extracted by the first branch to obtain a second component; Performing cross-attention calculation based on the high-frequency component and the second component to obtain a third component; Performing information gain on the third component based on the feature component extracted by the first branch to obtain a fourth component; Performing cross-attention calculation based on the low-frequency component and the fourth component to obtain a fifth component; Performing information gain on the fifth component based on the feature component extracted by the first branch to obtain a sixth component; A seventh component is obtained by performing cross-attention calculation based on the low-frequency component and the sixth component.
4. The remote sensing image change detection method according to claim 3, characterized in that: The decoder is specifically used for: Based on the first component and the third component extracted by each encoder branch, feature fusion is performed to obtain a first fusion feature; Perform feature fusion based on the fifth component and the seventh component extracted by each encoder branch to obtain a second fused feature; Perform feature fusion based on the fifth component and the seventh component extracted by each encoder branch to obtain a third fused feature; Based on the first component, the third component, the fifth component and the seventh component respectively extracted by each encoder branch, feature fusion is performed to obtain a fourth fusion feature; Perform feature fusion based on the first fusion feature and the fourth fusion feature to obtain a fifth fusion feature; Perform feature fusion based on the second fusion feature and the third fusion feature to obtain a sixth fusion feature; Change detection is performed based on the fifth fusion feature and the sixth fusion feature to obtain a change detection result.
5. The remote sensing image change detection method according to claim 4, characterized in that: The step of fusing the first component and the third component extracted by each encoder branch to obtain a first fused feature includes: Performing differential fusion on the first components extracted by each encoder branch to obtain a first differential fusion feature; Performing differential fusion on the third components extracted by each encoder branch to obtain a second differential fusion feature; The first differential fusion feature and the second differential fusion feature are scale-aligned to obtain a first fusion feature.
6. The remote sensing image change detection method according to claim 5, characterized in that: The step of performing scale alignment on the first differential fusion feature and the second differential fusion feature to obtain a first fusion feature includes: Performing upsampling processing on the second differential fusion feature to obtain a third differential fusion feature; The third differential fusion feature and the first differential fusion feature are concatenated to obtain a first fusion feature.
7. A remote sensing image change detection device, characterized in that: include: An acquisition module, used for acquiring two-phase remote sensing images to be detected; A detection module, used for inputting two-phase remote sensing images to be detected into a change detection model to obtain a change detection result output by the change detection model; wherein the change detection model is constructed based on an adjusted twin neural network; the adjusted twin neural network includes at least two encoder branches and one decoder; each encoder branch is used to extract a multi-scale semantic code of an input image by combining a first branch and a second branch thereof; the decoder is used to fuse and encode the multi-scale semantic codes extracted by each encoder branch under the guidance of high- and low-frequency information to obtain a change detection result; The first branch is a pre-trained visual basic model; the second branch is used to perform multi-stage semantic encoding on the input image, and the semantic encoding of each stage is guided by high-frequency and low-frequency information respectively.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the remote sensing image change detection method according to any one of claims 1 to 6 is implemented.
9. A storage medium, the storage medium being a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, the remote sensing image change detection method as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the remote sensing image change detection method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
High-resolution remote sensing image change detection method
CN113706482A
Deep learning-based remote sensing image cultivated land non-agricultural change detection method
CN115690591A
Twin change detection method based on information interaction and fusion
CN119206528A
Intelligent vibration digital twin systems and methods for industrial environments
US20210157312A1
Cited By
Remote sensing image semantic segmentation method fusing frequency domain modeling and lightweight linear attention
CN121170289A
Remote sensing image semantic segmentation method fusing frequency domain modeling and light linear attention
CN121170289B
Self-adaptive frequency collaborative remote sensing change detection method, device and system
CN121392574A