Building damage assessment method
By introducing multiple rounds of adaptive aptive global feature extraction and multi-scale wavelet fusion technology in the building damage assessment model, the problems of poor generalization and accuracy in the existing technology are solved, and more efficient and accurate building damage assessment is achieved.
Patent Information
- Application Number
- CN202510144071.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-10
AI Technical Summary
When facing complex post-disaster scenarios, existing building damage assessment technologies have problems such as poor generalization, poor accuracy and slow information processing speed, which are difficult to meet the needs of large-scale post-disaster building damage assessment.
A building damage assessment method is proposed, using the multi-round adaptive global feature extraction and multi-scale wavelet fusion technology of the building damage assessment model, and the evaluation accuracy and generalization ability of the model are improved through wavelet transformation and adaptive attention mechanism.
This method significantly improves the evaluation accuracy and generalization ability of the building damage assessment model, can more accurately identify the damage situation of the building, and perform well in complex scenarios to meet the needs of large-scale post-disaster building damage assessment.
Smart Images

Figure CN119942349A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image processing, in particular to the technical field of remote sensing image processing, and more particularly to a building damage assessment method. Background Art
[0002] Buildings, as indispensable spaces for human habitation and work, play a core role in urban and rural development. Rapid and accurate assessment of large-scale buildings is a key means to obtain disaster information and assist relevant personnel in formulating post-disaster rescue and reconstruction strategies. Since most buildings are located in densely populated areas, real-time assessment of their damage can predict the risk of collapse and assess the potential threat to surrounding residents. Therefore, in-depth research and development of rapid assessment technology for buildings, especially post-disaster building damage assessment technology, is particularly important and urgent.
[0003] Existing post-disaster building damage assessment technologies mainly use machine learning methods and convolutional neural networks to extract the appearance information of buildings for assessment. These methods perform well when the building outline is clear and significantly higher than the surrounding environment. However, the situation in reality is complex and changeable, especially in disaster-stricken areas. At the same time, the appearance and characteristics of buildings are affected by multiple factors such as region, scene, and lighting. Some buildings are difficult to distinguish because of their similar materials and colors to the background, while others are irregular in size and shape or are blocked or damaged, which increases the difficulty of extraction. In addition, the generalization ability of existing methods is limited. When there are obstructions, such as trees, other buildings or debris, these obstructions will interfere with the identification and assessment of buildings. Furthermore, information processing speed and large-scale data processing capabilities are also a bottleneck, making it difficult to quickly respond to needs in emergency disaster situations. In summary, existing building damage assessment technologies generally lack wide applicability and practicality as well as high accuracy, and are difficult to meet the needs of large-scale post-disaster building damage assessment. Summary of the invention
[0004] In view of the above problems, the present invention provides a building damage assessment method with improved generalization and assessment accuracy.
[0005] According to a first aspect of the present invention, there is provided a building damage assessment method, comprising:
[0006] The first image processing branch of the building damage assessment model is used to perform wavelet transform processing and multi-round adaptive global feature extraction on the pre-disaster images of the target building complex to obtain a multi-scale first feature atlas;
[0007] Using the first image processing branch to perform mask processing on the multi-scale first feature atlas, to obtain a first image segmentation result with mask information;
[0008] The second image processing branch of the building damage assessment model is used to perform wavelet transform processing, multi-round adaptive global feature extraction, and multi-round same-dimensional feature map fusion on the pre-disaster images and post-disaster images of the target building complex in parallel to obtain a multi-scale second feature map set.
[0009] Using the second image processing branch, convolution processing and activation processing are performed on the multi-scale second feature atlas to obtain a second image segmentation result;
[0010] The building damage assessment model is used to perform image fusion on the first image segmentation result and the second image segmentation result to obtain the damage assessment result of the target building complex.
[0011] According to an embodiment of the present invention, the first image processing branch of the building damage assessment model performs wavelet transform processing and multi-round adaptive global feature extraction on the pre-disaster image of the target building complex to obtain a multi-scale first feature atlas including:
[0012] The pre-disaster images of the target building complex are normalized and standardized using the building damage assessment model to obtain the pre-disaster images after preprocessing.
[0013] The multi-scale wavelet fusion module of the first image processing branch is used to perform a shunt wavelet transform process on the pre-processed pre-disaster image, and multiple initial wavelet transform feature maps are spliced and convolution-reduced to obtain a wavelet transform feature map;
[0014] The first image processing branch is used to divide the wavelet transform feature map into stripe regions in the vertical direction and the horizontal direction, and a neighborhood attention mechanism is applied to the wavelet transform feature map having multiple stripe regions to obtain a wavelet transform feature map having neighborhood information;
[0015] Multiple adaptive cross-attention modules of the first image processing branch are used to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set.
[0016] According to an embodiment of the present invention, the multi-scale wavelet fusion module using the first image processing branch performs a shunt wavelet transform process on the pre-processed pre-disaster image, and performs a splicing process and a convolution dimensionality reduction process on the obtained multiple initial wavelet transform feature maps, and the obtained wavelet transform feature map includes:
[0017] The pre-processed pre-disaster image is convolved using the feature map generation flow of the multi-scale wavelet fusion module, and the result of the convolution process is batch normalized to generate a first initial wavelet transform feature map;
[0018] Based on Haar wavelet transform, the image decomposition flow of the multi-scale wavelet fusion module is used to decompose the pre-disaster image after preprocessing into high-frequency components and low-frequency components, and multiple second initial wavelet transform feature maps are obtained;
[0019] The multi-scale wavelet fusion module is used to splice the first initial wavelet transform feature map and multiple second initial wavelet transform feature maps in the channel dimension, and the spliced feature map is subjected to convolution dimensionality reduction processing to obtain a wavelet transform feature map.
[0020] According to an embodiment of the present invention, the first image processing branch is used to divide the wavelet transform feature map into stripe areas in the vertical direction and the horizontal direction, and the wavelet transform feature map having multiple stripe areas is applied with the neighborhood attention mechanism to obtain the wavelet transform feature map having neighborhood information, including:
[0021] Using the first image processing branch, the wavelet transform feature map is divided into stripe regions in the vertical direction and the horizontal direction to obtain a wavelet transform feature map having a plurality of stripe regions;
[0022] Based on the preset window size information, each stripe region in the wavelet transform feature map having multiple stripe regions is divided into multiple cross-attention window regions by using the first image processing branch;
[0023] Based on the neighborhood attention mechanism, the first image processing branch is used to perform neighborhood feature fusion on all pixels in each cross-attention window area to obtain a wavelet transform feature map with neighborhood information.
[0024] According to an embodiment of the present invention, the above-mentioned neighborhood attention mechanism is based on the first image processing branch to perform neighborhood feature fusion on all pixels in each cross attention window area, and the wavelet transform feature map with neighborhood information is obtained, including:
[0025] Extracting local features of each pixel in each cross-attention window region using the first image processing branch to generate a local feature matrix for each pixel;
[0026] Using the first image processing branch, a learnable linear transformation is performed on the local feature matrix of each pixel to obtain a query vector, a key-value vector, and a value vector of each pixel;
[0027] Calculate the attention weight of each pixel using the query vector, the key vector and the value vector of each pixel, and embed the relative position deviation obtained by the convolution of the first image processing branch into the attention weight;
[0028] The first image processing branch is used to perform weighted summation of the attention weight and the local feature matrix of each pixel to obtain a wavelet transform feature map with neighborhood information.
[0029] According to an embodiment of the present invention, the above-mentioned use of multiple adaptive cross attention modules of the first image processing branch to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set includes:
[0030] Introduce the compression and excitation-based attention mechanism into multiple adaptive cross-attention modules, and set the hierarchical structure of each adaptive cross-attention module;
[0031] Using the first adaptive cross attention module to perform global average pooling on the wavelet transform feature map with neighborhood information, a first pooled feature map with context information is obtained;
[0032] Using the first adaptive cross attention module to perform full connection processing and nonlinear activation processing on the wavelet transform feature map with neighborhood information, a first activated feature map with channel attention weights is obtained;
[0033] The first adaptive cross-attention module is used to perform upsampling feature fusion based on the same channel dimension using the first pooled feature map with context information and the first activated feature map with channel attention weights to obtain the first attention feature map.
[0034] According to an embodiment of the present invention, the above-mentioned use of multiple adaptive cross attention modules of the first image processing branch to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set also includes:
[0035] Using the second adaptive cross attention module to perform global average pooling on the first attention feature map to obtain a second pooled feature map with context information;
[0036] Using the second adaptive cross attention module to perform full connection processing and nonlinear activation processing on the first attention feature map, a second activation feature map with channel attention weights is obtained;
[0037] The second adaptive cross-attention module is used to perform upsampling feature fusion based on the same channel dimension using the second pooled feature map with contextual information and the second activated feature map with channel attention weights to obtain the second attention feature map.
[0038] According to an embodiment of the present invention, the above-mentioned use of multiple adaptive cross attention modules of the first image processing branch to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set also includes:
[0039] Using the third adaptive cross attention module to perform global average pooling on the second attention feature map to obtain a third pooled feature map with context information;
[0040] The third adaptive cross-attention module is used to perform full connection processing and nonlinear activation processing on the second attention feature map to obtain a third activated feature map with channel attention weights.
[0041] According to an embodiment of the present invention, the above-mentioned using the first image processing branch to perform mask processing on the multi-scale first feature atlas to obtain the first image segmentation result with mask information includes:
[0042] Using the first image branch to upsample the feature maps in the multi-scale first feature map set, and performing splicing and fusion of the upsampled feature maps with the same channel dimension to obtain a spliced and fused first feature map;
[0043] The preset activation function of the first image branch is used to activate the first feature map after splicing and fusion, so as to obtain the first image segmentation result of the mask information.
[0044] According to an embodiment of the present invention, the second image processing branch of the building damage assessment model performs wavelet transform processing, multi-round adaptive global feature extraction, and multi-round same-dimensional feature map fusion on the pre-disaster image and the post-disaster image of the target building complex in parallel, and obtains a multi-scale second feature map set including:
[0045] Using different processing units of the second image processing branch to perform wavelet transform on the pre-disaster image and the post-disaster image respectively, and fusing the wavelet transform results of the pre-disaster image and the wavelet transform results of the post-disaster image with the same dimension to obtain a wavelet transform fusion feature map;
[0046] Using different processing units of the second image processing branch, respectively, multiple rounds of adaptive global features are performed on the wavelet transform results of the pre-disaster image and the wavelet transform results of the post-disaster image to obtain a plurality of pre-disaster global feature maps of different scales and a plurality of post-disaster global feature maps of different scales;
[0047] The pre-disaster global feature map and the post-disaster global feature map in each round are fused with features of the same dimension to obtain multiple global feature maps of different scales.
[0048] The building damage assessment method provided by the present invention can learn the robust features of buildings at different scales by performing multi-scale wavelet fusion on the images of the target building group, thereby improving the assessment accuracy and generalization ability of the building damage assessment model; at the same time, by applying a multi-round adaptive global attention mechanism to the images of the target building group, the key features of the target building group can be efficiently extracted, significantly enhancing the feature representation ability of the building damage assessment model. In addition, the building damage assessment method provided by the present invention further improves the overall performance and generalization ability of the building damage assessment model through deep image fusion technology, so that the building damage assessment model can perform building damage assessment in complex scenes and further improve the recognition of details of the target building group. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0050] Figure 1 is an application scenario diagram of a building damage assessment method according to an embodiment of the present invention;
[0051] Figure 2 is a flow chart of a building damage assessment method according to an embodiment of the present invention;
[0052] Figure 3 is a damage assessment architecture diagram of a target building complex according to an embodiment of the present invention;
[0053] Figure 4 is a data processing schematic diagram of a first image processing branch according to an embodiment of the present invention;
[0054] Figure 5 is a data processing schematic diagram of a second image processing branch according to an embodiment of the present invention;
[0055] Figure 6 is a structural diagram of a multi-scale wavelet fusion module according to an embodiment of the present invention;
[0056] Figure 7 is a structural diagram of an adaptive attention mechanism module according to an embodiment of the present invention;
[0057] Figure 8 is a structural diagram of a feedforward neural network with a gating mechanism according to an embodiment of the present invention;
[0058] Fig. 9 is a schematic diagram of training a building damage assessment model according to an embodiment of the present invention;
[0059] Fig.10is a schematic diagram of the result of post-disaster building damage assessment based on a mixed data set of xBD and xFBD according to an embodiment of the present invention;
[0060] Fig.11 is a schematic diagram of the result of building damage assessment in the Ida-BD data set according to an embodiment of the present invention;
[0061] Fig.12 A schematic diagram of a structure block diagram of a building damage assessment device according to an embodiment of the present invention is shown;
[0062] Fig.13 It is a block diagram of an electronic device suitable for implementing a building damage assessment method and a training method for a building damage assessment model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0063] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.
[0064] The terms used herein are only for describing specific embodiments, and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components. All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used here should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner. In the case of using a statement similar to "at least one of A, B, and C, etc.", it should generally be interpreted according to the meaning of the statement generally understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, C, etc.).
[0065] Building damage assessment has important application value in practice, especially post-disaster building damage assessment, which is of great significance for post-disaster rescue strategies and post-disaster reconstruction plans.
[0066] However, existing building damage assessment technologies have technical problems such as poor generalization and poor accuracy, and cannot meet the requirements of large-scale building damage assessment scenarios, especially in complex post-disaster scenarios. Therefore, it is necessary to provide a building damage assessment technology solution with high assessment accuracy, wide applicability and practicality.
[0067] An embodiment of the present invention provides a building damage assessment method, which applies the technical field of remote sensing image processing and post-disaster large-scale building damage assessment scenarios.
[0068] This paper innovatively proposes a building damage assessment method based on a damage assessment fusion network DAFN-FAAM (Damage Assessment Fusion Network with Adaptive Attention Mechanism) driven by an adaptive attention mechanism. The core of the DAFN-FAAM model lies in the Adaptive Attention Mechanism (AAM) module it introduces. The AAM module uses the adaptive attention mechanism to efficiently extract the key features of damaged buildings and significantly enhances the feature representation ability of the model. Although the traditional Transformer performs well in processing global features, it still has limitations in capturing local features and spatial information. To make up for this deficiency, the present invention further proposes a feedforward neural network (GMFFN) enhanced with a squeeze and excitation (SE) gating mechanism. By combining 3×3 deep convolution with SE attention gating, SEFFN significantly improves the overall performance and generalization ability of the model, making it perform well in processing complex scenes and detail recognition. In addition, in order to learn the robust features of buildings at different scales, the DAFN-FAAM model also introduces a multi-scale wavelet fusion (MWF) module. This module further improves the evaluation accuracy and generalization ability of the model by fusing multi-scale information. In order to verify the effectiveness and generalization ability of the DAFN-FAAM model, the present invention has been rigorously tested on large-scale disaster datasets xBD and xFBD. The test results show that the DAFN-FAAM model has achieved the most advanced performance level in terms of accuracy, and has shown significant advantages in processing complex scenes and detail recognition. In addition, in order to further verify the generalization ability of the model, the present invention also introduced the high-resolution satellite image dataset Ida-BD, and tested it with a hurricane in a certain place in 2021 as the center. The test results confirm the outstanding performance of the DAFN-FAAM model in high-precision damage assessment. The DAFN-FAAM framework has not only achieved significant improvements in evaluation accuracy, but also demonstrated excellent multi-scale feature fusion capabilities and cross-dataset migration performance. In practical applications, the framework has shown good adaptability and robustness, and can effectively cope with complex and diverse ground environments in remote sensing images. Therefore, the DAFN-FAAM framework has a high potential for promotion and application, and is expected to provide more powerful technical support and decision-making basis for emergency response and disaster management.
[0069] The application scenarios, specific technical solutions, training methods of building damage assessment models and assessment verification methods of the above-mentioned building damage assessment method provided by the present invention are described in detail below through specific embodiments.
[0070] Figure 1is an application scenario diagram of the building damage assessment method according to an embodiment of the present invention.
[0071] like Figure 1 As shown, the application scenario 100 according to this embodiment may include remote sensing image processing and post-disaster large-scale building damage assessment. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0072] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0073] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0074] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0075] It should be noted that the building damage assessment method provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the building damage assessment device provided in the embodiment of the present invention can generally be set in the server 105. The building damage assessment method provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the building damage assessment device provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0076] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0077] The following will be based on Figure 1 The scene described by Figure 2~Figure 8 The building damage assessment method of the disclosed embodiment is described in detail.
[0078] Figure 2 is a flow chart of a building damage assessment method according to an embodiment of the present invention.
[0079] like Figure 2 As shown, the building damage assessment method of this embodiment includes operations S210 to S250.
[0080] In operation S210, the pre-disaster image of the target building complex is subjected to wavelet transform processing and multi-round adaptive global feature extraction using the first image processing branch of the building damage assessment model to obtain a multi-scale first feature atlas.
[0081] The pre-disaster images and post-disaster images of the target building complex involved in the present invention may optionally be remote sensing images.
[0082] The above-mentioned building damage assessment model is constructed based on DAFN-FAAM. The building damage assessment model is divided into two image branches. The first image branch is used to utilize the pre-disaster images of the target building complex. The pre-disaster images are subjected to wavelet transform processing and multi-round adaptive global feature extraction. The multi-scale wavelet fusion module (MWF) of the DAFN-FAAM model is used for wavelet transform, and the adaptive attention module (AAM) of the DAFN-FAAM model is used for multi-round adaptive global feature extraction.
[0083] The first image processing branch of the above building damage assessment model includes four stages, namely, one stage of wavelet transform processing and three stages of adaptive global feature extraction.
[0084] In operation S220, a first image processing branch is used to perform mask processing on the multi-scale first feature atlas to obtain a first image segmentation result with mask information.
[0085] In operation S230, the second image processing branch of the building damage assessment model is used to perform wavelet transform processing, multi-round adaptive global feature extraction, and multi-round same-dimensional feature map fusion on the pre-disaster images and post-disaster images of the target building complex in parallel to obtain a multi-scale second feature map set.
[0086] Similar to the pre-disaster image processing process, the processing process of the second image processing branch for pre-disaster images and post-disaster images also includes one stage of wavelet transform processing and three stages of adaptive global feature extraction.
[0087] The second image processing branch of the building damage assessment model is used to process the pre-disaster images and post-disaster images of the target building complex in parallel. During the wavelet transform processing and multi-round adaptive global feature extraction, the pre-disaster images and post-disaster images are fused with multi-resolutions of the same scale before image processing is performed in the subsequent stage. For example, the wavelet transform results of the pre-disaster images are fused with the wavelet transform results of the post-disaster images for multi-resolution fusion and then processed subsequently.
[0088] In operation S240, a second image processing branch is used to perform convolution processing and activation processing on the multi-scale second feature atlas to obtain a second image segmentation result.
[0089] In operation S250, the first image segmentation result and the second image segmentation result are fused using a building damage assessment model to obtain a damage assessment result of the target building complex.
[0090] The damage assessment results of the above-mentioned target building complex include the background part of the image, the undamaged part, the slightly damaged part, the severely damaged part and the completely destroyed part.
[0091] The building damage assessment method provided by the present invention can learn the robust features of buildings at different scales by performing multi-scale wavelet fusion on the images of the target building group, thereby improving the assessment accuracy and generalization ability of the building damage assessment model; at the same time, by applying a multi-round adaptive global attention mechanism to the images of the target building group, the key features of the target building group can be efficiently extracted, significantly enhancing the feature representation ability of the building damage assessment model. In addition, the building damage assessment method provided by the present invention further improves the overall performance and generalization ability of the building damage assessment model through deep image fusion technology, so that the building damage assessment model can perform building damage assessment in complex scenes and further improve the recognition of details of the target building group.
[0092] The following is a specific embodiment and attached Figures 3-5 The building damage assessment method involved in the present invention is further described in detail.
[0093] Figure 3 4 is a diagram of a damage assessment framework for a target building complex according to an embodiment of the present invention.
[0094] Figure 4 is a data processing schematic diagram of a first image processing branch according to an embodiment of the present invention.
[0095] Figure 5is a data processing schematic diagram of a second image processing branch according to an embodiment of the present invention.
[0096] The satellite remote sensing images of the xBD and xFBD data sets are randomly cropped into blocks of 512×512 pixels, with the images having three RBG channels. A test set is determined from the cropped data set, and the building damage assessment method provided by the present invention is illustrated using the test set.
[0097] Based on the existing high-resolution semantic segmentation structure, the present invention proposes a building damage assessment method based on the deep learning network DAFN-FAAM. The data processing architecture of the above DAFN-FAAM is as follows: Figure 2 As shown in the figure. DAFN-FAAM is divided into two image processing branches, namely the first image processing branch and the second image processing branch, both of which are image segmentation networks. The first image processing branch uses a Transformer-based network to segment buildings in pre-disaster images and generates location masks of buildings. Pre-disaster images are randomly cropped into 512×512 images and fed to the first image processing branch for building segmentation.
[0098] The first image processing branch processes the pre-disaster image of the target building complex as follows: Figure 4As shown in the figure: First, the image is normalized and standardized. In the first stage, the 512×512 image is downsampled to a 128×128 feature map through MWF to provide a basis for the subsequent stages. After entering the second stage, the network first passes through four layers of AAM. The network not only maintains the high-resolution feature map of the first stage, but also generates a lower resolution (64×64) feature map through downsampling operations, forming two parallel feature map sets with different resolutions. The two feature map sets interact through a specific fusion strategy (for example, adding after restoring the resolution through upsampling), realizing preliminary multi-scale information fusion. The third stage also passes through four layers of AAM to further increase the resolution diversity of the feature map, generating three feature maps including the original high resolution (128×128), intermediate resolution (64×64) and lower resolution (32×32). These feature maps also interact through fusion strategies, allowing the network to capture richer contextual information. In the fourth stage, after only two layers of AAM, the first image processing branch finally generates feature maps of four resolutions: the highest resolution (128×128), the higher resolution (64×64), the intermediate resolution (32×32), and the lowest resolution (32×32). These feature maps are upsampled to the same size and deeply interact and fuse through the fusion mechanism of channel dimension splicing, making full use of information at different scales. Finally, through Softmax activation, accurate and detailed semantic segmentation results are generated through four times upsampling. Throughout the whole process, the first image processing branch always maintains the representation of high-resolution features.
[0099] The second image processing branch processes the pre-disaster images and post-disaster images of the target building complex as follows: Figure 5 As shown in the figure: Using two images before and after the disaster, the damage of the buildings after the disaster is classified through the Transformer-based fusion network (i.e., the second image processing branch). The second image processing branch has two secondary branches, each of which has the same structure as the second image processing branch. The difference is that in the first to fourth stages, all feature maps of the two secondary branches are obtained, and then each feature map is spliced according to the channel dimension, and then the channel dimension of the feature map is halved through a 3×3 convolution and fed to the next stage of the two secondary branches. The final feature map is subjected to a 3×3 convolution, and then a Softmax result is obtained. The result is combined with the mask generated by the first image processing branch, and after post-processing, the building damage is finally evaluated.
[0100] According to an embodiment of the present invention, the above-mentioned use of the first image processing branch of the building damage assessment model to perform wavelet transform processing and multi-round adaptive global feature extraction on the pre-disaster images of the target building complex to obtain a multi-scale first feature atlas includes: using the building damage assessment model to normalize and standardize the pre-disaster images of the target building complex to obtain pre-processed pre-disaster images; using the multi-scale wavelet fusion module of the first image processing branch to perform shunting wavelet transform processing on the pre-processed pre-disaster images, and performing splicing processing and convolution dimensionality reduction processing on the obtained multiple initial wavelet transform feature maps to obtain wavelet transform feature maps; using the first image processing branch to divide the wavelet transform feature map into stripe areas in the vertical and horizontal directions, and applying the neighborhood attention mechanism to the wavelet transform feature map with multiple stripe areas to obtain a wavelet transform feature map with neighborhood information; using multiple adaptive cross-attention modules of the first image processing branch to perform multi-round adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature atlas.
[0101] The following is a specific implementation method and combined with the attached Figures 6 and 7 The process of obtaining the multi-scale first feature atlas of the present invention is further described in detail.
[0102] Figure 6 4 is a structural diagram of a multi-scale wavelet fusion module according to an embodiment of the present invention.
[0103] Figure 7 is a structural diagram of an adaptive attention mechanism module according to an embodiment of the present invention.
[0104] Figure 8 is a structural diagram of a feedforward neural network with a gating mechanism according to an embodiment of the present invention.
[0105] The acquisition process of the above-mentioned multi-scale first feature atlas mainly involves wavelet transform and adaptive global feature extraction, which utilizes the multi-scale wavelet fusion (MWF) module and adaptive attention (AAM) module of DAFN-FAAM.
[0106] Since the scales of buildings in images of different disasters and regions are different, it is necessary to effectively capture the fine damage characteristics and background information of buildings at different scales. Therefore, a multi-scale wavelet fusion (MWF) module is introduced into the building damage assessment model. The structure of MWF is as follows: Figure 6 The MWF module uses wavelet transform to decompose the input image into high-frequency and low-frequency components, thereby capturing the building at different damage scales. This enables the model to capture key fine damage features and broader contextual information, such as the spatial relationship between the surrounding environment and the building, which is critical for accurate building damage assessment.
[0107] According to an embodiment of the present invention, the multi-scale wavelet fusion module of the first image processing branch is used to perform a shunt wavelet transform on the pre-processed pre-disaster image, and the obtained multiple initial wavelet transform feature maps are spliced and convolution-reduced to obtain a wavelet transform feature map, including: performing convolution on the pre-processed pre-disaster image using the feature map generation flow of the multi-scale wavelet fusion module, and performing batch normalization on the result of the convolution process to generate a first initial wavelet transform feature map; based on the Haar wavelet transform, the image decomposition flow of the multi-scale wavelet fusion module is used to decompose the pre-processed pre-disaster image into high-frequency components and low-frequency components to obtain multiple second initial wavelet transform feature maps; the first initial wavelet transform feature map and multiple second initial wavelet transform feature maps are spliced in the channel dimension using the multi-scale wavelet fusion module, and the spliced feature maps are convolution-reduced to obtain a wavelet transform feature map.
[0108] Although the first image processing branch and the second image processing branch are similar in structure, the pre-disaster and post-disaster images are transformed at two different scales.
[0109] First, the original image resolution is (512×512), and this image is passed through MWF to obtain a (128×128) feature map. The features are extracted through wavelet transform and embedded into the network encoder, as shown in Figure 5 As shown in the figure. The MWF module consists of two streams. The first stream operates on images of size (512×512) at the original resolution, using convolutional layers and batch normalization to generate feature maps. The second stream processes 4 feature maps of size (256×256) obtained by Haar wavelet transform. Haar wavelet decomposes the image into low-frequency components (YL) and high-frequency components (YH). The low-frequency component (YL) retains the overall structure and main features of the image, while the high-frequency component (YH) captures finer details and edge information. The feature maps in these two streams are then concatenated for further integration by splicing them together in the channel dimension. In order to ensure consistency with the number of channels introduced in the first stage of the second image processing branch, convolution is applied to the concatenated feature maps for dimensionality reduction, resulting in a final feature map size of (128×128). By embedding image features of different scales into the network, the MWF module can effectively capture the fine damage features of the building while enhancing broader contextual information.
[0110] According to an embodiment of the present invention, the above-mentioned use of the first image processing branch to divide the wavelet transform feature map into stripe areas in the vertical direction and the horizontal direction, and applying the neighborhood attention mechanism to the wavelet transform feature map with multiple stripe areas to obtain a wavelet transform feature map with neighborhood information includes: using the first image processing branch to divide the wavelet transform feature map into stripe areas in the vertical direction and the horizontal direction to obtain a wavelet transform feature map with multiple stripe areas; based on preset window size information, using the first image processing branch to divide each stripe area in the wavelet transform feature map with multiple stripe areas into multiple cross-attention window areas; based on the neighborhood attention mechanism, using the first image processing branch to perform neighborhood feature fusion on all pixels in each cross-attention window area to obtain a wavelet transform feature map with neighborhood information.
[0111] The division of the stripe area mentioned above helps capture the damage characteristics of the building and facilitates the subsequent classification task.
[0112] According to an embodiment of the present invention, the above-mentioned neighborhood attention mechanism is based on the use of the first image processing branch to perform neighborhood feature fusion on all pixels in each cross-attention window area to obtain a wavelet transform feature map with neighborhood information, including: using the first image processing branch to extract the local features of each pixel in each cross-attention window area to generate a local feature matrix for each pixel; using the first image processing branch to perform a learnable linear transformation on the local feature matrix of each pixel to obtain a query vector, a key-value vector and a value vector for each pixel; using the query vector, the key-value vector and the value vector of each pixel to calculate the attention weight of each pixel, and embed the relative position deviation obtained by the convolution of the first image processing branch into the attention weight; using the first image processing branch to perform weighted summation on the attention weight and the local feature matrix of each pixel to obtain a wavelet transform feature map with neighborhood information.
[0113] According to an embodiment of the present invention, the above-mentioned use of multiple adaptive cross-attention modules of the first image processing branch to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set includes: introducing a compression and excitation-based attention mechanism into multiple adaptive cross-attention modules, and setting a hierarchical structure of each adaptive cross-attention module; using the first adaptive cross-attention module to perform global average pooling processing on the wavelet transform feature map with neighborhood information to obtain a first pooled feature map with contextual information; using the first adaptive cross-attention module to perform full connection processing and nonlinear activation processing on the wavelet transform feature map with neighborhood information to obtain a first activated feature map with channel attention weights; using the first adaptive cross-attention module to perform upsampling feature fusion based on the same channel dimension on the first pooled feature map with contextual information and the first activated feature map with channel attention weights to obtain a first attention feature map.
[0114] According to an embodiment of the present invention, the above-mentioned use of multiple adaptive cross-attention modules of the first image processing branch to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set also includes: using the second adaptive cross-attention module to perform global average pooling processing on the first attention feature map to obtain a second pooling feature map with contextual information; using the second adaptive cross-attention module to perform full connection processing and nonlinear activation processing on the first attention feature map to obtain a second activation feature map with channel attention weights; using the second adaptive cross-attention module to perform upsampling feature fusion based on the same channel dimension on the second pooling feature map with contextual information and the second activation feature map with channel attention weights to obtain a second attention feature map.
[0115] According to an embodiment of the present invention, the above-mentioned use of multiple adaptive cross-attention modules of the first image processing branch to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain a multi-scale first feature map set also includes: using the third adaptive cross-attention module to perform global average pooling processing on the second attention feature map to obtain a third pooling feature map with contextual information; using the third adaptive cross-attention module to perform full connection processing and nonlinear activation processing on the second attention feature map to obtain a third activated feature map with channel attention weights.
[0116] According to an embodiment of the present invention, the above-mentioned use of the first image processing branch to perform mask processing on the multi-scale first feature atlas to obtain a first image segmentation result with mask information includes: using the first image branch to upsample the feature maps in the multi-scale first feature atlas, and splicing and fusing the upsampled feature maps with the same channel dimension to obtain the spliced and fused first feature maps; using the preset activation function of the first image branch to activate the spliced and fused first feature maps to obtain the first image segmentation result with mask information.
[0117] According to an embodiment of the present invention, the second image processing branch of the building damage assessment model is used to perform wavelet transform processing, multi-round adaptive global feature extraction and multi-round same-dimensional feature map fusion on the pre-disaster images and post-disaster images of the target building complex in parallel to obtain a multi-scale second feature map set, including: using different processing units of the second image processing branch to perform wavelet transform on the pre-disaster images and the post-disaster images respectively, and performing feature fusion of the same dimension on the wavelet transform results of the pre-disaster images and the wavelet transform results of the post-disaster images to obtain wavelet transform fusion feature maps; using different processing units of the second image processing branch to perform multi-round adaptive global features on the wavelet transform results of the pre-disaster images and the wavelet transform results of the post-disaster images respectively to obtain multiple pre-disaster global feature maps of different scales and multiple post-disaster global feature maps of different scales; performing feature fusion of the same dimension on the pre-disaster global feature maps and the post-disaster global feature maps in each round to obtain multiple global feature maps of different scales.
[0118] The above multi-scale first feature atlas provided by the present invention is further described in detail below through a specific implementation method. At the same time, since the second image processing branch is similar to the first image processing branch in structure, obtaining the multi-scale second feature atlas is similar to obtaining the multi-scale first feature atlas, and is combined and described here.
[0119] In order to enhance the robustness of the building damage assessment model, this paper proposes an innovative adaptive attention (AAM) module, such as Figure 7As shown. The AAM module combines the advantages of CNN and Transformer, introduces an adaptive cross-attention mechanism, reduces computational complexity, and achieves efficient global feature extraction. The AAM module adopts a cross-window-based self-attention mechanism, which can dynamically adjust the attention range and strike a balance between efficient computation and flexibility to adapt to different task requirements and input features. The process of processing feature maps is as follows: When processing the input feature map I∈R2C×H×W, the present invention adopts an innovative method, firstly dividing the feature map into stripe regions along the horizontal and vertical directions, the width and height of these stripes are set to h1 and w1 respectively, and ensuring that they are larger than the size of the local attention window k. In each h1×w1 stripe region, the present invention further divides the region into several k×k windows and applies the neighborhood attention mechanism therein. Specifically, for each pixel point coordinate (i, j) in the k×k window, the present invention extracts its local features to form a matrix Xp(i, j). The matrix is then transformed by a learnable transformation to generate the Query, Key, and Value matrices Q, K, and V. Next, combined with the relative position deviation B(i,j), the present invention calculates the attention weight Aij, where dk represents the dimension of Key. According to formula (1), the weighted sum of Aij and Vp(i,j) results in Z(i,j), which represents the fusion feature of pixel (i,j) and its neighborhood. After iterating all pixels in the stripe area, the present invention connects the output features Z(i,j) to obtain the feature map Z(i,j), which reflects the result of the local attention operation. To further enhance the feature representation, the present invention applies global attention in each h1×w1 stripe area. At this stage, the feature map within the window is reshaped into a two-dimensional matrix, where each row represents a feature channel and each column represents a pixel position within the window. Then, a standard self-attention mechanism is applied by performing a linear transformation to obtain new Query, Key, and Value matrices. By combining the relative position encoding BE(i,j), the attention weight is calculated and used to perform a weighted sum on the Value matrix, thereby obtaining a feature representation rich in global information. Finally, the weighted sum result is combined with the convolution position deviation Cij, as shown in formula (2), and reshaped into its original spatial dimension to obtain the global attention feature map of each stripe region. By connecting the global attention feature maps of all stripe regions, the present invention obtains the final feature map Yij. This feature map helps to capture the characteristics of building damage and provide a richer and more comprehensive representation for subsequent classification tasks.
[0120] (1),
[0121] (2).
[0122] In order to solve the limitations of Transformer in capturing local features and spatial information and improve the generalization ability of the model, this paper draws on the SE attention mechanism and proposes an innovative feedforward neural network (GMFFN). GMFFN introduces the SE mechanism to enhance the feature representation ability along the channel dimension. Figure 8 As shown in the figure, the present invention divides the original feature map into two parts along the channel dimension. The first part uses global average pooling to efficiently concentrate the input feature map and extract global context information. The second part processes the concentrated information through a fully connected layer and a nonlinear activation function (ReLU) to generate accurate channel attention weights. These weights reflect the importance of different channel features and achieve dynamic and accurate control of channel characteristics. Finally, the weights are multiplied element by element with the original input feature map to achieve fine-tuning of the channel features. The size of the final feature map is similar to the original input feature map. Figure 1 This design introduces a refined attention mechanism for the channel dimension while retaining the global information processing capability of Transformer. Therefore, GMFFN can effectively capture and utilize the potential relationships and differences between channel features, thereby improving the accuracy and efficiency of the model in building damage assessment.
[0123] Fig. 9 4 is a schematic diagram of training a building damage assessment model according to an embodiment of the present invention.
[0124] like Fig. 9 As shown in the figure, in order to obtain the above-mentioned building damage assessment model, it is necessary to use the satellite remote sensing images of the xBD and xFBD datasets to train the above-mentioned building damage assessment model. Due to the limitation of computer computing power, the image size is randomly cropped into 512×512 pixel blocks, and the image is RBG three channels. The dataset is divided into training set, validation set and test set, and verified on the Ida-BD dataset.
[0125] The present invention uses Focal Loss and Dice Loss functions to cope with the challenge of class imbalance in the data set to optimize the learning process of the model. The specific loss function is shown in formulas (3) to (5). Focal Loss is a loss function designed specifically to solve the problem of class imbalance. It adds an adjustment factor on the basis of traditional cross entropy loss, so that the model can pay more attention to samples that are difficult to distinguish. This adjustment factor is dynamic and gradually decreases to zero as the model's confidence in predicting the correct category increases. This means that during the training process, the model will pay less attention to samples that can be accurately classified, and more energy will be focused on samples that are difficult to correctly classify. Specifically, y represents the true label of the sample, and p represents the predicted probability of the model. This formula essentially measures the difference between the predicted result and the true label, and the final loss value is obtained by calculating 1 minus this difference. Therefore, when the predicted result is completely consistent with the true label, the loss value is zero; when the predicted result is completely inconsistent with the true label, the loss value reaches the maximum value, that is, 1:
[0126] (3),
[0127] (4),
[0128] (5).
[0129] During model training, the input data was first cropped to images of size 512×512. The experiments were conducted using the Pytorch 2.2.0 platform, with computing support provided by the NVIDIA RTX-3090 GPU. Various data augmentation strategies were used to enhance data diversity in the xBD and xFBD mixed datasets, including rotation, cropping, adding noise, and flipping. For gradient updates, the AdamW optimizer was selected because of its advantages in improving the training process. AdamW enhances the traditional Adam optimizer by incorporating a more sophisticated weight decay adjustment method, thereby promoting better generalization and faster convergence. The learning rate for the building localization (stage 1) and damage classification (stage 2) stages was consistently set to 1e-4. The model in stage 1 achieved the best performance at epoch 150, while the model in stage 2 achieved the best performance at epoch 300, utilizing the pre-trained weights in stage 1. For the Ida-BD dataset, the present invention did not incorporate it into the training process, but tested it separately. This approach ensures that the performance of the model of the present invention on high-resolution satellite imagery is unbiasedly validated.
[0130] Following the standard practice in building damage assessment research, we select precision and recall as the main metrics for evaluating model performance. This choice ensures that the evaluation method is consistent with the established benchmarks in the research community. In this paper, precision is defined as the ratio of true positives (TP) to the sum of true positives (TP) and false positives (FP), indicating the accuracy of positive predictions. Recall, also known as sensitivity, is calculated as the ratio of true positives (TP) to the sum of true positives (TP) and false negatives (FP), indicating the ability of the model to accurately identify all actual building damage instances. We utilize the building damage assessment metrics proposed in the xView2 Computer Vision Challenge, as shown in formula (6). These metrics include the weighted average F1 score (F1loc) for building segmentation and the unified average F1 score (F1cls) for damage classification, where F1Ci is the F1 score for each damage category and Ci is the i-th damage level. Specifically, C1 to C4 correspond to no damage, slight damage, severe damage, and damage, respectively. The definitions of F1loc and F1cls are shown in formulas (7) and (8).
[0131] (6),
[0132] (7),
[0133] (8).
[0134] In order to better illustrate the advantages of the trained DAFN-FAAM-based building damage assessment method, the following table is based on Table 1 and Table 2 and combined with the attached Fig.10 And attached Fig.11 To verify the effectiveness of the building damage assessment method provided by the present invention.
[0135] Fig.10 4 is a schematic diagram of the result of post-disaster building damage assessment based on a mixed data set of xBD and xFBD according to an embodiment of the present invention.
[0136] Fig.11 It is a schematic diagram of the results of building damage assessment in the Ida-BD dataset according to an embodiment of the present invention.
[0137] Table 1-Statistical table of comprehensive indicators for post-disaster building damage assessment
[0138] (xBD and XFBD datasets mixed)
[0139]
[0140] Table 2-Ida-BD dataset verification comprehensive index statistics
[0141]
[0142] Through Table 1 and Table 2 and Fig.10 and Fig.11 As shown, the building damage assessment method based on DAFN-FAAM of the present invention is far superior to the existing building damage assessment method in terms of actual display effect and verification index.
[0143] Based on the above-mentioned building damage assessment method, the present invention also provides a building damage assessment device. Fig.12 The device is described in detail.
[0144] Fig.12 The structure block diagram of the building damage assessment device according to an embodiment of the present invention is schematically shown.
[0145] like Fig.12 As shown, the building damage assessment device 1200 of this embodiment includes a first acquisition module 1210 , a mask processing module 1220 , a second acquisition module 1230 , a convolution and activation module 1240 , and an image fusion module 1250 .
[0146] The first acquisition module 1210 is used to use the first image processing branch of the building damage assessment model to perform wavelet transform processing and multi-round adaptive global feature extraction on the pre-disaster image of the target building complex to obtain a multi-scale first feature atlas; in one embodiment, the first acquisition module 1210 can be used to execute the operation S210 described above, which will not be repeated here.
[0147] The mask processing module 1220 is used to perform mask processing on the multi-scale first feature atlas using the first image processing branch to obtain a first image segmentation result with mask information. In one embodiment, the mask processing module 1220 can be used to perform the operation S220 described above, which will not be repeated here.
[0148] The second acquisition module 1230 is used to use the second image processing branch of the building damage assessment model to perform wavelet transform processing, multi-round adaptive global feature extraction and multi-round same-dimensional feature map fusion on the pre-disaster images and post-disaster images of the target building complex in parallel to obtain a multi-scale second feature map set; in one embodiment, the second acquisition module 1230 can be used to execute the operation S230 described above, which will not be repeated here.
[0149] The convolution and activation module 1240 is used to perform convolution and activation processing on the multi-scale second feature atlas using the second image processing branch to obtain a second image segmentation result. In one embodiment, the convolution and activation module 1240 can be used to perform the operation S240 described above, which will not be repeated here.
[0150] The image fusion module 1250 is used to use the building damage assessment model to perform image fusion on the first image segmentation result and the second image segmentation result to obtain a damage assessment result of the target building complex; in one embodiment, the image fusion module 1250 can be used to perform the operation S250 described above, which will not be repeated here.
[0151] According to an embodiment of the present invention, any multiple modules of the first acquisition module 1210, the mask processing module 1220, the second acquisition module 1230, the convolution and activation module 1240, and the image fusion module 1250 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the first acquisition module 1210, the mask processing module 1220, the second acquisition module 1230, the convolution and activation module 1240, and the image fusion module 1250 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or by any one of the three implementation methods of software, hardware, and firmware, or by a suitable combination of any of them. Alternatively, at least one of the first acquisition module 1210, the mask processing module 1220, the second acquisition module 1230, the convolution and activation module 1240, and the image fusion module 1250 can be at least partially implemented as a computer program module, and when the computer program module is executed, the corresponding function can be performed.
[0152] Fig.13 It is a block diagram of an electronic device suitable for implementing a building damage assessment method and a training method for a building damage assessment model according to an embodiment of the present invention.
[0153] like Fig.13 As shown, the electronic device 1300 according to an embodiment of the present invention includes a processor 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage part 1308 to a random access memory (RAM) 1303. The processor 1301 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1301 may also include an onboard memory for caching purposes. The processor 1301 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0154] In RAM 1303, various programs and data required for the operation of electronic device 1300 are stored. Processor 1301, ROM 1302 and RAM 1303 are connected to each other via bus 1304. Processor 1301 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 1302 and / or RAM 1303. It should be noted that the program can also be stored in one or more memories other than ROM 1302 and RAM 1303. Processor 1301 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.
[0155] According to an embodiment of the present invention, the electronic device 1300 may further include an input / output (I / O) interface 1305, which is also connected to the bus 1304. The electronic device 1300 may further include one or more of the following components connected to the input / output (I / O) interface 1305: an input portion 1306 including a keyboard, a mouse, etc.; an output portion 1307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1308 including a hard disk, etc.; and a communication portion 1309 including a network interface card such as a LAN card, a modem, etc. The communication portion 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the input / output (I / O) interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed, so that a computer program read therefrom is installed into the storage portion 1308 as needed.
[0156] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0157] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 1302 and / or RAM 1303 described above and / or one or more memories other than ROM 1302 and RAM 1303.
[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. It can be understood by those skilled in the art that the features recorded in the various embodiments of the present invention can be combined and / or combined in various ways, even if such a combination or combination is not explicitly recorded in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recorded in the various embodiments of the present invention can be combined and / or combined in various ways. All these combinations and / or combinations fall within the scope of the present invention.
[0159] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A building damage assessment method, characterized in that: The method comprises: The first image processing branch of the building damage assessment model is used to perform wavelet transform processing and multi-round adaptive global feature extraction on the pre-disaster images of the target building complex to obtain a multi-scale first feature atlas; Using the first image processing branch to perform mask processing on the multi-scale first feature atlas to obtain a first image segmentation result with mask information; Using the second image processing branch of the building damage assessment model, the pre-disaster images and post-disaster images of the target building complex are processed by wavelet transform, multi-round adaptive global feature extraction, and multi-round same-dimensional feature map fusion in parallel to obtain a multi-scale second feature map set; Using the second image processing branch to perform convolution processing and activation processing on the multi-scale second feature atlas to obtain a second image segmentation result; The building damage assessment model is used to perform image fusion on the first image segmentation result and the second image segmentation result to obtain a damage assessment result of the target building complex.
2. The method according to claim 1, characterized in that The first image processing branch of the building damage assessment model is used to perform wavelet transform processing and multi-round adaptive global feature extraction on the pre-disaster images of the target building complex, and the multi-scale first feature atlas is obtained, including: Using the building damage assessment model to normalize and standardize the pre-disaster image of the target building complex to obtain a pre-processed pre-disaster image; Using the multi-scale wavelet fusion module of the first image processing branch to perform a shunt wavelet transform process on the pre-processed pre-disaster image, and performing a splicing process and a convolution dimensionality reduction process on the obtained multiple initial wavelet transform feature maps to obtain a wavelet transform feature map; Using the first image processing branch to divide the wavelet transform feature map into stripe regions in the vertical direction and the horizontal direction, and applying a neighborhood attention mechanism to the wavelet transform feature map having multiple stripe regions to obtain a wavelet transform feature map having neighborhood information; The plurality of adaptive cross-attention modules of the first image processing branch are used to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain the multi-scale first feature map set.
3. The method according to claim 2, characterized in that The multi-scale wavelet fusion module of the first image processing branch is used to perform a shunt wavelet transform process on the pre-disaster image after pre-processing, and the obtained multiple initial wavelet transform feature maps are spliced and convolution-reduced to obtain a wavelet transform feature map including: Using the feature map generation flow of the multi-scale wavelet fusion module to perform convolution processing on the pre-disaster image after preprocessing, and performing batch normalization processing on the result of the convolution processing to generate a first initial wavelet transform feature map; Based on Haar wavelet transform, the pre-processed pre-disaster image is decomposed into high-frequency components and low-frequency components using the image decomposition flow of the multi-scale wavelet fusion module to obtain a plurality of second initial wavelet transform feature maps; The multi-scale wavelet fusion module is used to splice the first initial wavelet transform feature map and multiple second initial wavelet transform feature maps in the channel dimension, and the spliced feature map is subjected to convolution dimensionality reduction processing to obtain the wavelet transform feature map.
4. The method according to claim 2, characterized in that: The first image processing branch is used to divide the wavelet transform feature map into stripe regions in the vertical direction and the horizontal direction, and a neighborhood attention mechanism is applied to the wavelet transform feature map having multiple stripe regions to obtain a wavelet transform feature map having neighborhood information, including: Using the first image processing branch to divide the wavelet transform feature map into stripe regions in the vertical direction and the horizontal direction, so as to obtain a wavelet transform feature map having a plurality of stripe regions; Based on preset window size information, using the first image processing branch to divide each of the stripe regions in the wavelet transform feature map having the multiple stripe regions into a plurality of cross attention window regions; Based on the neighborhood attention mechanism, the first image processing branch is used to perform neighborhood feature fusion on all pixels in each cross-attention window area to obtain the wavelet transform feature map with neighborhood information.
5. The method according to claim 4, characterized in that Based on the neighborhood attention mechanism, the first image processing branch is used to perform neighborhood feature fusion on all pixels in each cross attention window area to obtain the wavelet transform feature map with neighborhood information, including: Extracting local features of each pixel in each cross-attention window region by using the first image processing branch, and generating a local feature matrix of each pixel; Using the first image processing branch, a learnable linear transformation is performed on the local feature matrix of each pixel to obtain a query vector, a key-value vector, and a value vector of each pixel; Calculating an attention weight of each pixel using the query vector, the key vector, and the value vector of each pixel, and embedding the relative position deviation obtained by the convolution of the first image processing branch into the attention weight; The first image processing branch is used to perform weighted summation on the attention weight and the local feature matrix of each pixel to obtain the wavelet transform feature map with neighborhood information.
6. The method according to claim 2, characterized in that The plurality of adaptive cross-attention modules of the first image processing branch are used to perform multiple rounds of adaptive global feature extraction on the wavelet transform feature map with neighborhood information to obtain the multi-scale first feature map set, including: Introducing a compression and excitation-based attention mechanism into a plurality of the adaptive cross attention modules, and setting a hierarchical structure of each of the adaptive cross attention modules; Using a first adaptive cross attention module to perform global average pooling processing on the wavelet transform feature map with neighborhood information to obtain a first pooled feature map with context information; Using the first adaptive cross attention module to perform full connection processing and nonlinear activation processing on the wavelet transform feature map with neighborhood information to obtain a first activated feature map with channel attention weights; The first adaptive cross-attention module uses the first pooled feature map with contextual information and the first activated feature map with channel attention weights to perform upsampling feature fusion based on the same channel dimension to obtain a first attention feature map.
7. The method according to claim 6, characterized in that Also includes: Using a second adaptive cross attention module to perform global average pooling on the first attention feature map to obtain a second pooled feature map with context information; Using the second adaptive cross attention module to perform full connection processing and nonlinear activation processing on the first attention feature map to obtain a second activation feature map with channel attention weights; The second adaptive cross-attention module uses the second pooled feature map with contextual information and the second activated feature map with channel attention weights to perform up-sampled feature fusion based on the same channel dimension to obtain a second attention feature map.
8. The method according to claim 7, characterized in that Also includes: Using a third adaptive cross attention module to perform global average pooling processing on the second attention feature map to obtain a third pooled feature map with context information; The third adaptive cross-attention module is used to perform full connection processing and nonlinear activation processing on the second attention feature map to obtain a third activated feature map with channel attention weights.
9. The method according to claim 1, characterized in that: Using the first image processing branch to perform mask processing on the multi-scale first feature atlas to obtain a first image segmentation result with mask information includes: Using the first image branch to upsample the feature maps in the multi-scale first feature map set, and performing splicing and fusion of the upsampled feature maps with the same channel dimension to obtain a spliced and fused first feature map; The preset activation function of the first image branch is used to activate the first feature map after splicing and fusion, so as to obtain the first image segmentation result of the mask information.
10. The method according to claim 1, characterized in that The second image processing branch of the building damage assessment model is used to perform wavelet transform processing, multi-round adaptive global feature extraction and multi-round same-dimensional feature map fusion on the pre-disaster image and the post-disaster image of the target building complex in parallel, and the multi-scale second feature map set is obtained, including: Using different processing units of the second image processing branch to perform wavelet transform on the pre-disaster image and the post-disaster image respectively, and performing feature fusion of the same dimension on the wavelet transform results of the pre-disaster image and the wavelet transform results of the post-disaster image to obtain a wavelet transform fusion feature map; Using different processing units of the second image processing branch, respectively, perform multiple rounds of adaptive global features on the wavelet transform results of the pre-disaster image and the wavelet transform results of the post-disaster image to obtain a plurality of pre-disaster global feature maps of different scales and a plurality of post-disaster global feature maps of different scales; The pre-disaster global feature map and the post-disaster global feature map in each round are fused with features of the same dimension to obtain multiple global feature maps of different scales.
Citation Information
Patent Citations
Post-disaster building damage assessment method, device and equipment based on remote sensing image
CN118314471A