Disaster response-oriented change detection method based on dual-time-phase characteristic interaction

By adopting a change detection method based on two-time phase feature interaction in building damage assessment after disasters, combined with the ConvMamba encoder and multi-branch expansion convolution feature enhancement module, the difficulties of feature extraction and complex background recognition in the prior art are solved, and higher detection accuracy and more reliable disaster response support are achieved.

CN120198431AInactive Publication Date: 2025-06-24NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510682020.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract small building features and identify difficult samples in complex backgrounds in post-disaster building damage assessment, resulting in low detection accuracy and difficulty in supporting disaster response and humanitarian rescue.

Method used

The change detection method based on dual-time phase feature interaction is adopted, and multi-scale features are extracted through the weight-sharing ConvMamba encoder, combined with the multi-branch expansion convolution feature enhancement module and the dual-time phase feature interaction module of grouped Mamba, and finally decode and damage assessment are performed in the multi-stage decoder.

Benefits of technology

It improves the accuracy of identifying damage levels of buildings after disasters, enhances the detection capabilities of small buildings and complex backgrounds, effectively solves the problems of mis-checking and missed inspections, and provides reliable decision-making support for disaster response and humanitarian rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198431A_ABST
    Figure CN120198431A_ABST
Patent Text Reader

Abstract

The invention provides a disaster response-oriented change detection method based on dual-time-phase characteristic interaction, which comprises the following steps of: 1, acquiring building remote sensing image data before and after a disaster, and preprocessing the building remote sensing image data; 2, sending the preprocessed building remote sensing image data into a weight sharing ConvMama encoder to extract multi-scale features of images before and after a disaster; 3, the multi-scale features are sent to a feature enhancement module based on multi-branch expansion convolution, and enhanced features are obtained; step 4, the enhanced features are sent to a double-time-phase feature interaction module based on grouping Mamba, and difference features are obtained; and step 5, sending the difference characteristics into a multi-level decoder for decoding, and outputting a building damage grading graph to complete damage evaluation. According to the method disclosed by the invention, accurate damage evaluation can be well realized on the building under various complex background influences, and the problems of false detection, missing detection and the like are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of electronics, communication, and information engineering, and particularly relates to a change detection method based on dual-temporal feature interaction for disaster response. Background Art

[0002] In recent decades, various natural disasters have occurred frequently around the world. Earthquakes, floods, hurricanes, and other extremely destructive natural disasters not only cause huge property losses but also pose a great threat to human life safety. Therefore, in addition to improving the detection and early warning capabilities before a disaster occurs, it is also necessary to quickly obtain disaster information after a disaster occurs. When a disaster occurs, rapid and accurate building damage assessment is crucial for humanitarian relief and disaster response. However, after a disaster, the communication and transportation infrastructure conditions are limited, and it is difficult to carry out the traditional building damage assessment method based on ground surveys. The rapid development of remote sensing technology has significantly improved the accuracy and efficiency of change detection technology. As a sub-task of change detection, building damage assessment has also benefited significantly. High-spatial-resolution remote sensing images can accurately reflect the ground surface and quickly provide large-area observation data to support building damage assessment.

[0003] In recent years, deep learning technology has achieved remarkable results in tasks such as remote sensing image change detection and semantic segmentation, providing new solutions for building damage assessment. However, the building damage assessment task still faces many challenges. For example, damaged buildings often only occupy a small area in the image and are easily overlooked; complex ground surface environments (such as shadows, clouds, occlusions, etc.) may lead to false detections or missed detections; in addition, the damage patterns caused by different disaster types to buildings vary greatly, posing higher requirements for the generalization ability of the model. Therefore, how to improve the detection ability of the model for small target damage areas in diverse scenarios and enhance its robustness to complex backgrounds is an important research direction at present. Summary of the Invention

[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a change detection method based on dual-temporal feature interaction for disaster response aiming at the deficiencies of the prior art, enhance the accuracy of post-disaster building damage assessment, solve the problems that traditional change detection methods cannot effectively extract the features of small buildings and the recognition accuracy of difficult samples is low, and provide reliable decision-making support for disaster response and humanitarian relief.

[0005] The method of the present invention includes the following steps: Step 1, collect remote sensing image data of buildings before and after a disaster, and preprocess the remote sensing image data of buildings. The preprocessing includes flipping, rotating, and cropping; Step 2: Feed the preprocessed building remote sensing image data into a weight - shared ConvMamba (Convolutional Mamba) encoder to extract multi - scale features of pre - disaster and post - disaster images; Step 3: Feed the multi - scale features into a feature enhancement module based on multi - branch dilated convolutions to obtain enhanced features; Step 4: Feed the enhanced features into a dual - temporal feature interaction module based on grouped Mamba to fully interact the dual - temporal features and obtain difference features; Step 5: Feed the difference features into a multi - level decoder for decoding, and output a building damage grading map through a detector to complete the damage assessment.

[0006] Step 2 includes: Based on the Vision Mamba architecture, build a weight - shared ConvMamba encoder. The ConvMamba encoder has four stages. In the first stage, the input preprocessed building remote sensing image data first passes through a stem layer (dry layer) for preliminary feature processing and dimensionality reduction, and then passes through a Convolutional Visual State Space Block (CVSSB) to extract the global and local spatial features of pre - disaster and post - disaster remote sensing images. The operations performed in the second, third, and fourth stages are the same, that is: first downsample the input data, and then use the Convolutional Visual State Space Block (CVSSB) to extract the global and local spatial features of pre - disaster and post - disaster remote sensing images. Finally, the Convolutional Visual State Space Block (CVSSB) will fuse the global and local spatial features through a concatenation operation to obtain the final features of the ConvMamba encoder and ; where is the feature extracted by the pre - disaster image through the i - th stage of the ConvMamba encoder, is the feature extracted by the post - disaster image through the i - th stage of the ConvMamba encoder.

[0007] In Step 2, the following operations are performed in the Convolutional Visual State Space Block (CVSSB): The input feature is first split into two branches along the channel dimension, namely the convolutional branch and the State Space Module (SSM) branch. Denote the inputs of the two branches as and . The convolutional branch consists of two 3×3 convolutions and a ReLU function, and is used to extract local spatial features ; In the SSM branch, the input features are divided into two information flows after passing through a linear embedding layer and layer normalization. One of the information flows passes through a depth convolution layer (DW convolution, Depthwise Convolution) and the SiLU activation function, and then through the 2D Selective Scan module SS2D (2D Selective Scan). The output of the 2D Selective Scan module SS2D passes through a linear layer to obtain the initial global spatial features. The other information flow passes through a linear layer and the SiLU activation function to obtain the gating features. Then, the outputs of the convolution branch and the SSM branch are element-wise multiplied to obtain the finally output global spatial features. ; Then, the local spatial features and the global spatial features are fused through a concatenation operation to obtain the fused features. The fused features are subjected to a channel shuffle operation, and finally fused with the original features through a residual connection to obtain the final output. and The formula is expressed as: (1), (2), (3), (4), (5), Among them, represents the final output of the Convolutional Vision State Space Block CVSSB; ReLU(·) represents the ReLU activation function, BN(·) represents the batch normalization layer, Conv3×3 represents the 3×3 convolution operation, LN(·) represents the layer normalization, SiLU(·) represents the SiLU activation function, DWConv(·) represents the depth convolution, and Lin(·) represents the linear layer; represents element-wise multiplication, represents element-wise addition; Concat(·) represents the concatenation operation, and Shuffle(·) represents the channel shuffle operation.

[0008] In step 2, the SS2D module performs the following operations: Given the input data (the input data is the features output after passing through the depth convolution layer and the SiLU activation function), the SS2D module first flattens the input data into sequences along four directions, and the four directions refer to from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right. Then, the sequences in the four directions are respectively input into the Selective State Space Module for parallel processing, and finally the obtained sequences are integrated into a new feature map.

[0009] In step 3, the feature enhancement module includes 4 branches, namely the first branch, the second branch, the third branch and the fourth branch; the input of each branch (the input feature here is actually and , for the convenience of expression in the formula of this module, the input feature is represented as F in ) first passes through a 1×1 convolution to adjust the number of channels; the output of the first branch after passing through the 1×1 convolution is directly added to the outputs of the other three branches; the second branch performs two cascaded dilated convolution operations; the third branch sequentially performs 1×3, 3×1 and a dilated convolution operation; the fourth branch performs a 3×3 convolution operation, and the formula is expressed as: (6), (7), (8), (9), (10), Among them, , , and respectively represent the standard convolution operations with convolution kernel sizes of 1×1, 1×3, 3×1 and 3×3; and respectively represent the dilated convolution operation with a dilation rate of 3 and the dilated convolution operation with a dilation rate of 5; and respectively represent the input feature and the output feature of the feature enhancement module; , , , respectively represent the output features of the first branch, the second branch, the third branch and the fourth branch.

[0010] In step 4, the dual-temporal feature interaction module performs the following operations: cross-stitches the enhanced features obtained in step 3 along the channel dimension, and the formula is: (11), Among them, represents the mixed feature, represents the feature tensor of the (k + 1)-th channel in the mixed feature, represents the -th channel tensor of; represents the -th channel tensor of; , respectively refer to the features of the pre-disaster image and the post-disaster image processed by the feature enhancement module; Set the number of channels of the finally output mixed features to 2C, then is expressed as: (12), Set the number of channels of the mixed features to 2C, then the number of channels of the features of the pre-disaster image and the post-disaster image before mixing are both C, respectively represent the feature of the first channel of the pre-disaster image, the feature of the first channel of the post-disaster image, the feature of the C-th channel of the pre-disaster image, and the feature of the C-th channel of the post-disaster image; Divide the mixed features into four groups along the channel dimension, which will generate four branches. Each branch gets a sub-feature map, and a total of four sub-feature maps are obtained. Then use four Visual State Space Blocks (VSS) to perform feature extraction on the four sub-feature maps simultaneously, and then splice the outputs of the four branches. For , perform subtraction operations, and then further extract the change information through convolution operations. Then send the change information into the multiplication branch to fuse with the features processed by the VSS block, and finally obtain the difference features of the remote sensing image. The formula is expressed as: (13), (14), (15), Among them, represents the grouped Mamba operation, represents the features output after the grouped Mamba operation, represents the output features of the dual-temporal feature interaction module.

[0011] Step 5 includes: sending the difference features obtained by the dual-temporal feature interaction module into a multi-level decoder. The multi-level decoder has four stages. In each stage, the multi-level decoder first uses the Visual State Space Block (VSS) to model the global spatial context information of the current input features, and then inputs the globally modeled features into the fusion module. The fusion module simultaneously receives the features obtained after the processing of the adjacent stages, and fuses the globally modeled features with the features obtained after the processing of the adjacent stages to obtain the fused features; then input the fused features into the linear layer and the upsampling layer, and gradually restore the difference feature map to the size of the original input image. Finally, obtain the result of damage classification through a detector; In the fusion module, first, the features output after processing by the visual state space block VSS are input into a convolutional block to further enhance the feature expression ability and restore the detailed information. Then, the features processed in adjacent stages are input into the fusion module and added to the features processed by the convolutional block to obtain the first fused feature. The first fused feature is continuously input into the convolutional block for processing to obtain the second fused feature. Finally, the first fused feature is connected to the second fused feature through a residual connection, and the final feature fusion is completed through an adder. The convolutional block in the fusion module consists of a 3×3 convolution, a batch normalization layer, and a ReLU activation function. Cross-entropy loss is used to optimize the model training, and the formula is: (16), where K represents the number of classification categories, and K is 4 in the damage assessment task. represents the cross-entropy loss of the building damage assessment task, N represents the number of samples. represents the ground truth value of the k-th category. represents the output of the building damage classification network for the k-th category.

[0012] In step 5, Lovasz-softmax loss is introduced, and the final loss function is expressed as: (17), where, , represent the Lovasz-softmax loss of the building damage assessment task and the total loss of the building damage assessment task, respectively.

[0013] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method.

[0014] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, it executes the steps of the method.

[0015] Beneficial effects: 1. The present invention proposes a new building damage assessment method, which combines the local modeling ability of CNN and the global modeling ability of Mamba using CVSSB, effectively extracts and integrates the global and local information of the image, and improves the recognition accuracy of the damage degree of post-disaster buildings.

[0016] 2. The present invention designs a building feature enhancement module, which enriches the features of small buildings through multi-branch dilated convolutions, enhances the network's detection ability for small building targets, and solves the problem of loss of small building feature details during the feature extraction process.

[0017] 3. The present invention proposes a dual-temporal feature interaction module, which mixes the channel dimensions of dual-temporal image features and uses the grouped Mamba method to achieve full interaction of dual-temporal image features, extracts more accurate change features, and enhances the recognition ability for difficult samples.

[0018] 4. Compared with other damage assessment methods, the method proposed by the present invention can accurately assess the damage of buildings under various complex backgrounds, effectively solve problems such as false detection and missed detection, and has positive significance for humanitarian rescue and disaster response. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is the overall architecture diagram of the building damage assessment method proposed by the present invention.

[0020] Figure 2 is the specific flowchart of CVSSB (Convolutional Visual State Space Block).

[0021] Figure 3 is the specific structure diagram of SS2D (2D Selective Scan Module).

[0022] Figure 4 is the specific structure diagram of the feature enhancement module.

[0023] Figure 5 is the specific structure diagram of the feature interaction module.

[0024] Figure 6 is the specific structure diagram of the decoder block.

[0025] Figure 7 is the schematic diagram of the comparative experiment on the xBD dataset.

[0026] Figure 8 is the structure diagram of the computer device. DETAILED DESCRIPTION OF THE INVENTION

[0027] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0028] For damaged buildings, they often occupy only a small area in the image and are extremely easy to be ignored; complex surface environments (such as shadows, clouds, obstacles, etc.) may lead to false detections or missed detections; the damage forms caused by different disaster types to buildings vary greatly; existing models have poor recognition accuracy for difficult categories, etc. In the embodiments of the present invention, a remote sensing image change detection method based on dual-temporal feature interaction for disaster response is designed. In this method, first, the remote sensing image data of buildings before and after the disaster are collected, and preprocessing such as flipping, rotating, and cropping is performed on them; secondly, the preprocessed image set is sent into the designed encoder with shared weights to extract multi-scale features of the pre-disaster and post-disaster images; thirdly, these multi-scale features are sent into a feature enhancement module based on multi-branch dilated convolution to obtain richer building detail features; then, these enhanced features are sent into a dual-temporal feature interaction module based on grouped Mamba, and after fully interacting the dual-temporal features, accurate difference features are obtained. Finally, the difference features are sent into a multi-level decoder for decoding, and a building damage grading map is output through a detector to complete the damage assessment. The method designed by the present invention can combine the global spatial features and local spatial features of the image to obtain more comprehensive building features, and enhance the network's detection ability for small buildings through the feature enhancement module. In addition, the method designed by the present invention can fully interact the building features before and after the disaster, effectively improving the network's recognition ability for difficult categories. Therefore, the present invention can better solve some problems existing in the aforementioned related technologies.

[0029] Embodiments of the present invention provide a change detection method based on dual-temporal feature interaction for disaster response, as Figure 1 shown, the overall framework of the method consists of an encoder with shared weights, a feature enhancement module, a dual-temporal feature interaction module based on grouped Mamba, and a multi-level Mamba decoder.

[0030] For the sub-task of building damage assessment in this change detection, local spatial features play a crucial role in the detection of change features (difference features). However, most of the existing deep learning-based building damage assessment methods enhance the model's ability to extract global spatial features by designing different attention mechanisms or scanning mechanisms, while ignoring the role of local spatial features. Therefore, based on the VMamba architecture, the present invention proposes a ConvMamba encoder with shared weights, as Figure 1 shown, Figure 1Among them, P1, P2, P3, and P4 respectively represent the pre-disaster and post-disaster image feature pairs processed by four feature enhancement modules. The encoder has four stages. In the first stage, the input first passes through a stem layer for preliminary feature processing and dimensionality reduction to better extract features in subsequent stages, and then passes through a CVSSB to extract the global and local spatial features of the pre-disaster and post-disaster remote sensing images. In the subsequent three stages, the input data is first downsampled, and then the CVSSB is used to extract the global and local spatial features of the pre-disaster and post-disaster remote sensing images. Finally, the CVSSB will fuse the global and local spatial features through a concatenation operation to obtain the final features of this stage and 。

[0031] The structure of the CVSSB is as Figure 2 shown. The input feature is first split into two branches along the channel dimension, namely the convolutional branch and the SSM branch. The inputs of the two branches are respectively denoted as and . The convolutional branch consists of two 3×3 convolutions and the ReLU function, which is used to extract local spatial features . In the SSM branch, the input feature is divided into two information flows after passing through a linear embedding layer and layer normalization. One information flow passes through a depthwise separable convolutional layer and the SiLU activation function, and then passes through the SS2D module. The output of the SS2D module passes through a linear layer to obtain the initial global spatial feature , and the other information flow passes through a linear layer and the SiLU activation function to obtain the gating feature . Then, the outputs of the two branches are multiplied element-wise to obtain the final output global spatial feature, that is . Then, the local spatial feature and the global spatial feature are fused through a concatenation operation. The fused feature is subjected to a channel shuffle operation to avoid information loss between channels. Finally, it is fused with the original feature through a residual connection to obtain the final output and . This process is formulated as: (1), (2), (3), (4), (5), where represents the input of the convolutional branch, represents the input of the SSM branch, Represents the final output of CVSSB. ReLU(·) represents the ReLU activation function, BN(·) represents the batch normalization layer, Conv3×3 represents the 3×3 convolution operation, LN(·) represents the layer normalization, SiLU(·) represents the SiLU activation function, DWConv(·) represents the depthwise separable convolution, and Lin(·) represents the linear layer. Represents element-wise multiplication, Represents element-wise addition. Concat(·) represents the concatenation operation, and Shuffle(·) represents the channel shuffle operation.

[0032] SS2D (2D Selective Scan Module) is a core module in the CVSSB designed in the present invention, and its specific structure is as Figure 3 shown. Given the input data, SS2D first flattens the input into sequences along four directions, namely from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right, then inputs them into S6 for parallel processing respectively, and finally integrates the obtained sequences into a new feature map. Through the 2D scanning mechanism, all pixels can integrate the information of other pixels from different directions.

[0033] High-resolution optical remote sensing images are vulnerable to natural environments and imaging conditions, and due to the diversity and complexity of ground objects, problems such as the loss of small target feature details are likely to occur in change detection tasks such as building damage assessment. The encoder has limited ability to extract building features. Therefore, in order to capture richer small target semantic features, the present invention designs a feature enhancement module based on multi-branch dilated convolution. Its structure is as Figure 4 shown.

[0034] The input of each branch of the feature enhancement module designed in the present invention is first adjusted in the number of channels through a 1×1 convolution for subsequent processing. The output of the first branch is directly added to the outputs of the other three branches after passing through the 1×1 convolution. The design of this branch can well retain the features of the original image and avoid losing key features. The second branch performs two cascaded dilated convolution operations. The dilated convolution increases the receptive field, enabling the extracted features to retain more context information. The third branch sequentially performs 1×3, 3×1, and a dilated convolution operation. By decomposing the 3×3 convolution, the directional features of the image (such as horizontal and vertical features) can be further mined. The fourth branch performs a conventional 3×3 convolution operation to extract the local spatial features of the image and enhance the expression ability of information such as edges and textures. The above process can be formulated as: (6), (7), (8), (9), (10), Among them, , , and respectively represent standard convolution operations with convolution kernel sizes of 1×1, 1×3, 3×1, and 3×3; and respectively represent dilated convolution operations with dilation rates of 3 and 5; and respectively represent the input feature and output feature of the feature enhancement module; , , , respectively represent the output features of the four branches. Concat represents the feature concatenation operation.

[0035] The feature enhancement module designed in the present invention enables the network to learn richer context space features through a multi-branch dilated convolution structure, improving the network's feature extraction ability for small buildings. For some small-scale buildings in disasters, effective detection can be achieved.

[0036] For the building damage assessment task, it is crucial to fully interact the dual-temporal features of remote sensing images before and after the disaster and obtain accurate differential features. However, due to the complexity of high-resolution optical remote sensing images, simply subtracting or concatenating the features of dual-temporal images before and after the disaster cannot effectively extract the changed features. Therefore, a dual-temporal feature interaction module based on grouped Mamba is proposed in this paper.

[0037] The structure of the feature interaction module is as shown in Figure 5 . This module performs spatio-temporal modeling on dual-temporal images by mixing dual-temporal features in the channel dimension. Specifically, for the dual-temporal features enhanced by the feature enhancement module, they are first cross-concatenated along the channel dimension. Through this processing method, the network can more efficiently fuse and correlate the features of the two time phases, helping to capture more fine-grained change information and avoiding the problem that some change details may be lost when directly using the difference operation. This process is formulated as: (11), Among them, represents the mixed feature, represents the feature tensor of the (k + 1)-th channel in the mixed feature, represents the -th channel tensor of Represents the tensor of the -th channel; Assume that the number of channels of the finally output mixed feature is 2C, then it is expressed as: (12), By mixing the image features before and after the change in the channel dimension, the adjacent channels of the mixed feature respectively contain the semantic features of the bi-temporal images. Therefore, this paper designs a grouped Mamba method to process the mixed feature. Specifically, the mixed feature is divided into 4 groups along the channel dimension, and then 4 VSSs are used to extract features from 4 sub-feature maps simultaneously, and then the outputs of the four branches are concatenated. In order to enhance the change information in the mixed feature, the features of the bi-temporal remote sensing images are subtracted, and then the change information is further extracted through a convolution operation, and then the change information is sent to the multiplication branch to fuse with the features processed by VSS, and finally the difference features of the remote sensing image are obtained. The above process can be formulated as: (13), (14), (15), where, represents the grouped Mamba operation, represents the feature output after the grouped Mamba operation, represents the output feature of the feature interaction module.

[0038] After that, the difference features obtained by the feature interaction module are sent into a multi-level decoder. The multi-level decoder is mainly responsible for gradually restoring the multi-scale features output by the feature interaction module to the original image size. Similar to the encoder, the decoder also has four stages.

[0039] Figure 6 gives a simplified architecture of the decoder. Specifically, first use the VSS block to model the global spatial context information of the current input feature, then fuse it with the difference feature map from the adjacent stage through a fusion module, and finally gradually restore the feature map to the original image size through an upsampling operation. The specific architecture of the fusion module is also given in Figure 6 .

[0040] Since the post-disaster building damage assessment task can be regarded as a semantic change detection task, the cross-entropy loss can also be used to optimize the model training. Therefore, the cross-entropy loss is formulated as: (16), wherein, represents the cross-entropy loss for the building damage assessment task, N represents the number of samples, represents the ground truth value of the k-th class, represents the output of the building damage classification network for the k-th class. In addition, the present invention also introduces the Lovasz-softmax loss to alleviate the problem of sample number imbalance between changed and unchanged pixels. Therefore, the final loss function can be formulated as: (17), wherein, , respectively represent the Lovasz-softmax loss for the building damage assessment task and the total loss of this task.

[0041] To verify the effectiveness of the building damage assessment method proposed by the present invention, the present invention conducted a comparative experiment on the currently largest remote sensing dataset xBD for building damage assessment with several latest methods. The experimental results are as Figure 7 shown, where white represents no damage, blue represents slight damage, orange represents severe damage, and red represents complete damage. The experimental results show that compared with other methods, the method proposed by the present invention can extract more detailed building contours, and can better detect the damage categories of dense buildings, while other methods have misdetections to varying degrees. For small buildings, the method proposed by the present invention can successfully identify their damage categories, while other methods have all missed detections. Therefore, the above experiments prove the effectiveness and superiority of the method proposed by the present invention in building damage assessment.

[0042] The present invention also provides a computer device, and its structural schematic diagram is as Figure 8 shown. The computer device includes: a processor, a memory, and a computer program stored on the memory and executable on the processor. The processor implements the above-mentioned change detection method based on dual-temporal feature interaction for disaster response when executing the program. The processor, the memory, and the network interface are interconnected through a bus and complete communication. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media, etc., and is used to provide an operating environment and store experimental data. The network interface of the computer device is used to communicate with external terminals.

[0043] The present invention provides a change detection method based on dual-temporal feature interaction for disaster response. There are many methods and approaches to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

Claims

1. A change detection method based on dual-temporal feature interaction for disaster response, characterized in that Including the following steps: Step 1: Collect the remote sensing image data of the building before and after the disaster, and preprocess the remote sensing image data of the building. The preprocessing includes flipping, rotating, and cropping; Step 2: Feed the preprocessed remote sensing image data of the building into the weight - shared ConvMamba encoder to extract the multi - scale features of the pre - disaster and post - disaster images; Step 3: Feed the multi - scale features into the feature enhancement module based on multi - branch dilated convolution to obtain the enhanced features; Step 4: Feed the enhanced features into the dual - temporal feature interaction module based on grouped Mamba, and obtain the difference features after fully interacting the dual - temporal features; Step 5: Feed the difference features into the multi - level decoder for decoding, and output the building damage grading map through a detector to complete the damage assessment.

2. The method according to claim 1, wherein Step 2 includes: Based on the Vision Mamba architecture, a weight - shared ConvMamba encoder is established. The ConvMamba encoder has four stages. In the first stage, the pre - processed building remote - sensing image data is first passed through a stem layer for preliminary feature processing and dimensionality reduction, and then passed through a Convolutional Vision State - Space Block (CVSSB) to extract the global and local spatial features of the pre - and post - disaster remote - sensing images. The operations performed in the second, third, and fourth stages are the same, which are: first, downsample the input data, then use the Convolutional Vision State - Space Block (CVSSB) to extract the global and local spatial features of the pre - and post - disaster remote - sensing images. Finally, the Convolutional Vision State - Space Block (CVSSB) will fuse the global and local spatial features through a concatenation operation to obtain the final features of the ConvMamba encoder. and ; where is the feature extracted by the pre - disaster image through the i - th stage of the ConvMamba encoder, is the feature extracted by the post - disaster image through the i - th stage of the ConvMamba encoder.

3. The method according to claim 2, wherein In step 2, the following operations are performed in the Convolutional Visual State Space Block (CVSSB): the input features are first split into two branches along the channel dimension, namely the convolutional branch and the State Space Model (SSM) branch. Denote the inputs of the two branches as and respectively. The convolutional branch consists of two 3×3 convolutions and a ReLU function, which is used to extract local spatial features ; In the SSM branch, the input features are divided into two information streams after passing through a linear embedding layer and layer normalization. One of the information streams passes through a depth convolution layer and the SiLU activation function, and then through the 2D selective scanning module SS2D. The output of the 2D selective scanning module SS2D passes through a linear layer to obtain the initial global spatial features , and the other information stream passes through a linear layer and the SiLU activation function to obtain the gating features . Then, the outputs of the convolution branch and the SSM branch are multiplied element-wise to obtain the final output global spatial features ; Then, the local spatial features and the global spatial features are fused through a concatenation operation to obtain the fused features. The fused features are subjected to a channel shuffle operation, and finally fused with the original features through a residual connection to obtain the final output and , which is expressed by the formula: (1), (2), (3), (4), (5), Among them, represents the final output of the Convolutional Vision State Space Block CVSSB; ReLU(·) represents the ReLU activation function, BN(·) represents the batch normalization layer, Conv3×3 represents the 3×3 convolution operation, LN(·) represents the layer normalization, SiLU(·) represents the SiLU activation function, DWConv(·) represents the depth convolution, and Lin(·) represents the linear layer; represents element-wise multiplication, represents element-wise addition; Concat(·) represents the concatenation operation, and Shuffle(·) represents the channel shuffle operation.

4. The method according to claim 3, wherein In Step 2, the SS2D module performs the following operations: Given the input data, the SS2D module first flattens the input data into sequences along four directions, which are from top - left to bottom - right, from bottom - right to top - left, from top - right to bottom - left, and from bottom - left to top - right. Then, the sequences in the four directions are respectively input into the selective state - space model for parallel processing, and finally the obtained sequences are integrated into a new feature map.

5. The method according to claim 4, characterized in that, In Step 3, the feature enhancement module includes 4 branches, namely the first branch, the second branch, the third branch, and the fourth branch; The input of each branch first adjusts the number of channels through a 1×1 convolution; The output of the first branch after the 1×1 convolution is directly added to the outputs of the other three branches; The second branch performs two cascaded dilated convolution operations; The third branch sequentially performs 1×3, 3×1, and a dilated convolution operation; The fourth branch performs a 3×3 convolution operation, and the formula is expressed as: (6), (7), (8), (9), (10), Among them, , , and respectively represent standard convolution operations with convolution kernel sizes of 1×1, 1×3, 3×1, and 3×3; and respectively represent dilated convolution operations with dilation rates of 3 and 5; and respectively represent the input feature and the output feature of the feature enhancement module; , , , respectively represent the output features of the first branch, the second branch, the third branch, and the fourth branch.

6. The method according to claim 5, wherein In Step 4, the dual - temporal feature interaction module performs the following operations: Cross - splice the enhanced features obtained in Step 3 along the channel dimension, and the formula is: (11), Among them, represents a mixed feature, represents the feature tensor of the (k + 1)-th channel in the mixed feature, represents the -th channel tensor of; represents the -th channel tensor; and respectively refer to the features of the pre-disaster image and the post-disaster image after being processed by the feature enhancement module; Set the number of mixed feature channels for the final output to 2C, then It is expressed as: (12), Set the number of channels of the mixed feature to 2C, then the number of feature channels of the pre-disaster image and the post-disaster image before mixing is both C. They respectively represent the feature of the first channel of the pre-disaster image, the feature of the first channel of the post-disaster image, the feature of the C-th channel of the pre-disaster image, and the feature of the C-th channel of the post-disaster image. The mixed features are divided into four groups along the channel dimension, resulting in four branches. Each branch obtains a sub-feature map, and a total of four sub-feature maps are obtained. Then, four visual state space blocks (VSS) are used to extract features from the four sub-feature maps simultaneously. After that, the outputs of the four branches are concatenated, and subtraction operations are performed on , . Then, convolution operations are further used to extract the changing information. Subsequently, the changing information is sent to the multiplication branch to be fused with the features processed by the VSS blocks, and finally, the differential features of the remote sensing image are obtained. The formula is expressed as: (13), (14), (15), Among them, represents the grouped Mamba operation, represents the features output after the grouped Mamba operation, represents the output features of the dual-temporal feature interaction module.

7. The method according to claim 6, wherein Step 5 includes: Feed the difference features obtained by the dual - temporal feature interaction module into the multi - level decoder. The multi - level decoder has four stages. In each stage, the multi - level decoder first uses the visual state - space block VSS to model the global spatial context information of the current input features, and then inputs the globally modeled features into the fusion module. The fusion module simultaneously receives the features obtained after the processing of the adjacent stages, and fuses the globally modeled features with the features obtained after the processing of the adjacent stages to obtain the fused features; Then, the fused features are input into the linear layer and the up - sampling layer to gradually restore the difference feature map to the size of the original input image, and finally the result of damage classification is obtained through a detector; In the fusion module, first input the features output after the processing of the visual state - space block VSS into a convolution block to further enhance the feature expression ability and restore the detail information. Then, input the features processed by the adjacent stages into the fusion module and add them to the features processed by the convolution block to obtain the first fused feature; Input the first fused feature into the convolution block for further processing to obtain the second fused feature. Finally, the first fused feature is connected to the second fused feature through a residual connection, and the final feature fusion is completed through an adder; The convolutional block in the fusion module consists of a 3×3 convolution, a batch normalization layer, and a ReLU activation function; Cross-entropy loss is used to optimize model training, and the formula is: (16), where K represents the number of categories for classification, represents the cross-entropy loss for the building damage assessment task, N represents the number of samples, represents the ground truth of the k-th category, represents the output of the building damage classification network for the k-th category.

8. The method according to claim 7, characterized in that, In step 5, Lovasz-softmax loss is introduced, and the final loss function is expressed as: (17), Among them, , respectively represent the Lovasz-softmax loss of the building damage assessment task and the total loss of the building damage assessment task.

9. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, Stores a computer program or instruction, and when the computer program or instruction runs on a computer, it executes the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Remote sensing image building change detection method based on multi-scale attention

    CN118298305A

  • Dual-branch multi-scale dynamic local convolution attention method based on remote sensing change detection

    CN118736416A

  • Remote sensing image semantic change detection method and device based on Mamba model

    CN119580258A

  • Remote sensing image building damage assessment method combining global and local clues

    CN119600431A

Cited By

  • Remote sensing image building extraction method and system based on visual Mama model

    CN120599504A

  • A Forest Fire Severity Inversion Method Based on Dual-Temporal Feature Fusion and Boundary Constraints

    CN122574662A