Airport target damage assessment method, device and equipment based on multi-source image fusion
By using multi-source image fusion technology, combining optical and radar images, high-precision positioning and damage level assessment of airport target facilities have been achieved, solving the problems of low efficiency and poor accuracy in traditional methods and meeting the needs for rapid and accurate airport damage assessment.
Patent Information
- Application Number
- CN202511700683.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-02-24
AI Technical Summary
Traditional airport damage assessment methods rely on manual surveys or single-modal remote sensing images, which are inefficient and have poor accuracy, making it difficult to meet the needs of rapid and accurate airport post-disaster assessment.
An airport target damage assessment method based on multi-source image fusion is adopted. By acquiring multimodal data (optical and radar images) before damage and optical images after damage, and combining them with a pre-trained airport damage assessment model, feature extraction and fusion processing are performed to generate segmentation mask maps and damage level mask maps, thereby realizing fine-grained segmentation and damage level assessment of airport target facilities.
It improves the positioning accuracy and damage assessment accuracy of airport target facilities, enabling precise damage assessment under adverse conditions such as cloud cover and fog, and meeting the needs of refined assessment under complex airport structures.
Smart Images

Figure CN121564483A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and in particular to a method, apparatus and equipment for airport target damage assessment based on multi-source image fusion. Background Technology
[0002] As transportation hubs and strategic facilities, airports are vulnerable to damage in emergencies or natural disasters. Damage assessment is of great significance for emergency response and post-disaster recovery. Traditional damage assessment methods rely on manual surveys or single-modal remote sensing image analysis, which have problems such as low efficiency and poor accuracy. Summary of the Invention
[0003] This application provides a method, apparatus, and equipment for airport target damage assessment based on multi-source image fusion.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a method for airport target damage assessment based on multi-source image fusion. The method includes: in response to airport damage, acquiring multiple first optical images and multiple radar images of the airport before the damage, and multiple second optical images of the airport after the damage; based on a pre-constructed airport damage assessment model, obtaining a segmentation mask map of the airport by performing feature extraction and fusion processing on the multiple first optical images and multiple radar images; the segmentation mask map is used to characterize the location and type information of each target facility in the airport; based on the airport damage assessment model, obtaining the comparison results of each target facility before and after the damage by comparing the multiple first optical images and multiple second optical images; based on the airport damage assessment model, performing damage level assessment processing on the comparison results of each target facility before and after the damage and the airport segmentation mask map to obtain a damage level mask map of the airport; the damage level mask map is used to characterize the degree of damage to each target facility.
[0005] This application provides an airport target damage assessment device based on multi-source image fusion, comprising: an acquisition module, configured to acquire multiple first optical images and multiple radar images of the airport before damage, and multiple second optical images of the airport after damage, in response to damage to the airport; a first processing module, configured to obtain a segmentation mask map of the airport by performing feature extraction and fusion processing on the multiple first optical images and multiple radar images based on a pre-constructed airport damage assessment model; the segmentation mask map is used to characterize the location and type information of each target facility in the airport; a second processing module, configured to obtain the comparison results of each target facility before and after damage by comparing each target facility in the multiple first optical images and multiple second optical images based on the airport damage assessment model; and a third processing module, configured to perform damage level assessment processing on the comparison results of each target facility before and after damage and the segmentation mask map of the airport based on the airport damage assessment model, to obtain a damage level mask map of the airport; the damage level mask map is used to characterize the degree of damage to each target facility.
[0006] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the steps in any of the above methods.
[0007] In this embodiment, by acquiring pre-damage multimodal data (multiple first optical images and multiple radar images) and post-damage optical data (multiple second optical images), and combining them with a pre-trained airport damage assessment model, fine-grained segmentation and damage level assessment of target facilities in the airport are achieved. First, by extracting and fusing features from the pre-damage multimodal data, the complementarity of optical and radar images can be utilized to improve the positioning accuracy of target facilities and effectively address adverse conditions such as cloud and fog obstruction. Then, by comparing the pre-damage and post-damage optical images, changes in the target facilities over time can be dynamically captured, identifying these changes and providing data for subsequent damage level assessment. Finally, by fusing the segmentation mask image and the comparison results, precise classification of the damage level of each target facility is achieved, meeting the needs of refined assessment under complex airport structures and improving the accuracy and practicality of damage assessment. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the first process of an airport target damage assessment method based on multi-source image fusion provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the airport damage assessment model provided in the embodiments of this application; Figure 3This is a schematic diagram of the second process of an airport target damage assessment method based on multi-source image fusion provided in an embodiment of this application; Figure 4 This is a schematic diagram of the third process of an airport target damage assessment method based on multi-source image fusion provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an airport target damage assessment device based on multi-source image fusion provided in an embodiment of this application.
[0009] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] This application provides a method for airport target damage assessment based on multi-source image fusion, such as... Figure 1 As shown, the method includes steps S100 to S130: Step S100: In response to damage to the airport, acquire multiple first optical images and multiple radar images of the airport before the damage, and multiple second optical images of the airport after the damage.
[0012] In some implementations, the airport may be damaged by an emergency, such as a natural disaster or war damage.
[0013] It is understood that in this application, the airport needs to acquire multiple first optical images and multiple radar images before damage, where the optical images and radar images are data of different modes, that is, multimodal image data is acquired before damage.
[0014] Optical imaging is a technology that acquires, processes, and transmits images using optical principles. Its core lies in utilizing physical phenomena such as the propagation, reflection, refraction, and interference of light to capture light through optical devices, converting it into electrical or digital signals, and thus forming an image. These optical devices can include, but are not limited to, cameras, surveillance equipment, and infrared thermal imagers.
[0015] Radar imagery is generated by a radar system that transmits radio waves and receives the echo signals scattered by a target, ultimately producing an image that reflects the spatial characteristics of the target. Radar systems can include, but are not limited to, synthetic aperture radar (SAR) satellites, interferometric SAR satellites, and lidar.
[0016] In some implementations, the multiple first optical images may be optical images of the airport taken from different angles before the airport was damaged, and / or optical images of the airport taken at different times.
[0017] In some implementations, multiple radar images may be radar images taken of the airport from different angles before the airport was damaged, and / or radar images taken of the airport at different times.
[0018] In some implementations, the multiple first optical images may be optical images of the airport taken from different angles after the airport has been damaged, and / or optical images of the airport taken at different times.
[0019] Step S110: Based on the pre-built airport damage assessment model, a segmentation mask map of the airport is obtained by extracting and fusing features from multiple first optical images and multiple radar images; the segmentation mask map is used to characterize the location and type information of each target facility in the airport.
[0020] Here, the airport damage assessment model is a systematic analytical framework used to quantify the extent of damage to target facilities in an airport after a sudden event.
[0021] In some implementations, the airport damage assessment model's processing stages include a first stage and a second stage. The first stage is used to obtain a segmentation mask image of the airport, and the processing of this first stage may include shallow feature extraction, deep feature extraction, and segmentation mask image generation. The second stage may include a pre-disaster and post-disaster feature comparison process and a damage level assessment process. In this application, by employing a two-stage network architecture to separate target localization and damage assessment tasks, interference between target localization and damage assessment tasks is avoided, enhancing the robustness of the airport damage assessment model.
[0022] In some implementations, the processing module of the airport damage assessment model includes a shallow feature extraction and cross-modal fusion module, a deep feature extraction and cross-modal fusion module, an airport target segmentation decoder module, a cross-temporal feature fusion module, and a damage level assessment module. The shallow feature extraction and cross-modal fusion module, the deep feature extraction and cross-modal fusion module, and the airport target segmentation decoder module correspond to the first stage, while the cross-temporal feature fusion module and the damage level assessment module correspond to the second processing stage.
[0023] In this application, in the first processing stage, feature extraction processing is performed by inputting pre-disaster multimodal image data (multiple first optical images and multiple radar images) to extract shallow and deep features of the multimodal image data, and a segmentation mask map is determined by fusing the shallow and deep features.
[0024] In some implementations, the segmentation mask map labels the location and type of various target facilities in the airport, such as runways, aprons, hangars, and fuel depots, in pixels, where each pixel value represents a different category, thereby achieving fine segmentation of the airport's internal structure.
[0025] Step S120: Based on the airport damage assessment model, compare each target facility in multiple first optical images and multiple second optical images to obtain the comparison results of each target facility before and after damage.
[0026] In some implementations, multiple first optical images before the damage and multiple second optical images after the disaster are input into a cross-temporal feature fusion module. The cross-temporal feature fusion module compares the state of the same target facility in different temporal phases and outputs the comparison results.
[0027] For example, the cross-temporal feature fusion module compares the runway area identified in pre-damage images with targets in the same location in post-disaster images to determine whether cracks, collapses, or complete destruction have occurred. By calculating pixel-level differences and combining them with semantic segmentation results, it outputs the changes in each target facility, such as area reduction or shape alteration. This comparison method not only improves identification accuracy but also provides a basis for subsequent damage level assessment.
[0028] In some implementations, shallow features and deep semantic features of the same target facility at different time phases are obtained; based on the shallow features and deep semantic features at different time phases, the comparison result of the target facility is determined.
[0029] For example, shallow features of the same target facility under different time phases are compared to obtain a first feature comparison result; deep semantic features of the same target facility under different time phases are compared to obtain a second feature comparison result; and the comparison result is determined based on the first feature comparison result and the second feature comparison result.
[0030] For example, shallow features and deep semantic features of the same target facility at the same time phase are fused to obtain the first fused feature; the first fused features at different time phases are compared to determine the comparison result.
[0031] In some implementations, the second processing stage includes a first sub-processing stage and a second sub-processing stage, wherein the first sub-processing stage is used to determine the comparison results of each target facility before and after the damage, and the first sub-processing stage and the first processing stage can operate in parallel.
[0032] Step S130: Based on the airport damage assessment model, damage level assessment is performed on the comparison results of each target facility before and after damage and the segmentation mask map of the airport to obtain the damage level mask map of the airport; the damage level mask map is used to characterize the degree of damage of each target facility.
[0033] Here, a damage level mask is a two-dimensional image used to characterize the degree of damage to various target facilities in an airport, where each pixel value represents the corresponding damage level (e.g., undamaged, slightly damaged, severely damaged, completely damaged). This damage level mask can provide a quantitative basis for post-disaster recovery decisions.
[0034] In some implementations, the second processing stage includes a second sub-processing stage, wherein the second sub-processing stage is used to determine a damage level mask map.
[0035] In some implementations, the damage level of the target facility is determined based on the comparison results and a preset damage level strategy; the damage level of the target facility is then marked on a segmentation mask to obtain a damage level mask. The damage level of the target facility may include four levels: no damage, minor damage, severe damage, and complete destruction.
[0036] In this embodiment, by acquiring pre-damage multimodal data (multiple first optical images and multiple radar images) and post-damage optical data (multiple second optical images), and combining them with a pre-trained airport damage assessment model, fine-grained segmentation and damage level assessment of target facilities in the airport are achieved. First, by extracting and fusing features from the pre-damage multimodal data, the complementarity of optical and radar images can be utilized to improve the positioning accuracy of target facilities and effectively address adverse conditions such as cloud and fog obstruction. Then, by comparing the pre-damage and post-damage optical images, changes in the target facilities over time can be dynamically captured, identifying these changes and providing data for subsequent damage level assessment. Finally, by fusing the segmentation mask image and the comparison results, precise classification of the damage level of each target facility is achieved, meeting the needs of refined assessment under complex airport structures and improving the accuracy and practicality of damage assessment.
[0037] In some embodiments, step S110 includes steps S111 to S113: Step S111: Using a dual-branch convolutional neural network architecture, feature extraction processing is performed on multiple first optical images and multiple radar images to determine the first shallow features of each target facility in the multiple first optical images and the second shallow features of each target facility in the multiple radar images; the first shallow features and the second shallow features are used to characterize the attribute information of each target facility. Here, feature extraction refers to the process of extracting key features from the original image that are helpful for target recognition and classification. In this application, a dual-branch convolutional neural network is used to extract shallow features from optical images and radar images respectively, and cross-modal fusion is performed through an attention mechanism.
[0038] In some implementations, the dual-branch convolutional neural network includes a first-branch convolutional neural network and a second-branch convolutional neural network. The first-branch convolutional neural network performs feature extraction processing on multiple first optical images, outputting first shallow features; the second-branch convolutional neural network performs feature extraction processing on multiple radar images, outputting second shallow features. Alternatively, the first-branch convolutional neural network performs feature extraction processing on multiple radar images, outputting first shallow features; the second-branch convolutional neural network performs feature extraction processing on multiple first optical images, outputting second shallow features.
[0039] Understandably, a dual-branch convolutional neural network (DN) architecture can divide the input data into two independent processing paths, processing optical and radar images separately. This preserves the unique characteristics of each modality and avoids information loss. For example, in optical images, the first branch of the DN architecture can extract geometric features such as the runway's edge contour and the rectangular structure of the apron; while in radar images, the second branch can capture physical features such as the surface roughness of buildings and the intensity of metal reflection. Through this dual-path data processing approach, shallow features (first and second shallow features) of target facilities in different modalities can be extracted independently, providing high-quality foundational data for subsequent cross-modal fusion. Therefore, the design of the DN architecture directly affects the quality and diversity of shallow features and is one of the key prerequisites for determining a high-precision segmentation mask.
[0040] The first shallow feature layer consists of low-level visual features extracted from optical images, typically including local information such as edges, corners, and textures. These features are highly effective at describing the basic morphology of target facilities, such as the straight edges of a runway or the regular paving texture of an apron. The second shallow feature layer, derived from radar imagery, possesses stronger penetration capabilities and the ability to detect obstructed areas. For example, synthetic aperture radar imagery can detect the outline of target facilities hidden under clouds or vegetation. By extracting both the first and second shallow features, a more comprehensive depiction of the spatial distribution and physical state of target facilities can be achieved.
[0041] In some implementations, the attribute information of the target facilities may include, but is not limited to, one or more of the following: shape, texture, and outline. The shape represents the overall appearance characteristics of each target facility, such as the straightness of the runway and the cuboid structure of the hangar. The texture reflects the detailed features of the surface of each target facility, such as the brick paving pattern of the apron and the distribution of runway cracks. The outline is a continuous line of the boundary of each target facility, which can be used to accurately locate the boundary position of each target facility.
[0042] Step S112: Perform feature fusion processing on the first shallow feature and the second shallow feature to determine the shallow fusion feature of each target facility; Here, feature fusion processing refers to fusing shallow features from different modalities to generate more representative comprehensive features. The feature fusion process typically includes feature alignment, weighted summation, and attention mechanisms. Shallow fused features are comprehensive features generated based on features extracted by a dual-branch convolutional neural network architecture, using attention mechanisms and weighted summation. Shallow fused features contain both high-contrast details from optical images and structural information from radar images.
[0043] In some implementations, a combination of global average pooling and an attention mechanism is used to fuse the first and second shallow features. The attention mechanism dynamically adjusts the weights based on the importance of each channel, enhancing important features while suppressing irrelevant or interfering features. Finally, the attention results from each channel are weighted to generate the shallow fused feature.
[0044] Shallow fusion features are new feature vectors generated by fusing shallow features from optical and radar images. They can simultaneously reflect the information advantages of both optical and radar images. For example, in cloud-covered scenes, the penetrating power of radar images can be effectively used to supplement the features of missing areas in optical images. By fusing the first and second shallow features, a more accurate target description can be generated.
[0045] It should be noted that by performing feature fusion processing on the first and second shallow features, the modal differences between optical and radar images can be eliminated, thereby enabling the subsequent airport target segmentation decoder module to better understand the overall state of the target facility and thus improve the accuracy of calculating the segmentation mask.
[0046] Step S113: Determine the segmentation mask based on the first shallow feature, the second shallow feature, and the shallow fusion feature corresponding to all target facilities in the airport.
[0047] In some implementations, firstly, the first and second shallow features are input into the deep feature extraction and cross-modal fusion module to determine the deep semantic features of the target facility; then, the deep semantic features and shallow fused features are input into the airport target segmentation decoder module to determine the segmentation mask.
[0048] In this embodiment, firstly, shallow features from optical and radar images are extracted using a dual-branch convolutional neural network architecture, effectively capturing basic attribute information of the target facility, such as shape, texture, and contour. Then, by fusing shallow features from different modalities, the airport damage assessment model's ability to perceive the target facility is enhanced, improving its segmentation accuracy. Finally, a comprehensive model is built upon all shallow features and fused shallow features of the target facility to generate a more accurate segmentation mask, thus laying a data foundation for subsequent damage assessment of the target facility.
[0049] In some embodiments, step S113 includes steps S1131 and S1132: Step S1131: Based on the attention calculation of the first shallow features and the second shallow features corresponding to all target facilities respectively, determine the deep semantic features corresponding to all target facilities respectively; The first shallow features refer to the low-level features of optical images extracted by convolutional neural networks (CNNs), typically the outputs of the first few layers of backbone networks such as ResNet. These first shallow features can capture local details in optical images, such as the straight edges of a runway or the rectangular structure of an apron.
[0050] The second shallow feature is a low-level feature extracted from another modality of data (radar imagery) corresponding to the first shallow feature. The low-level feature has a similar function to the first shallow feature and can provide complementary information to the first shallow feature. For example, when optical images are obscured by clouds or fog, the penetrating power of radar images can supplement the missing edge information in the optical images.
[0051] Attention calculation is used to measure the correlation between different input features and dynamically assign weights. This application employs a combination of self-attention and cross-attention to interactively model shallow features within the same modality and shallow features across modalities, respectively. By using a combination of self-attention and cross-attention, the consistency between features within a modality and the synergistic relationship between features across modalities can be enhanced.
[0052] Deep semantic features are high-dimensional feature representations generated after attention mechanisms and global semantic enhancement. They not only preserve the geometric information of the original image but also integrate cross-modal contextual information. These deep semantic features generated by attention mechanisms and global semantic enhancement can more accurately represent the category and spatial distribution of target facilities, thus providing a more reliable basis for subsequent segmentation tasks.
[0053] In some implementations, the first shallow features and the second shallow features are input into the deep feature extraction and cross-modal fusion module, and the deep semantic features are determined by the deep feature extraction and cross-modal fusion module.
[0054] In some implementations, deep semantic features may include location and type information of the target facility. Location information refers to the specific coordinates and / or area of the target facility within the airport, such as the start and end points of a runway; location information is crucial for locating damaged areas. Type information is used to classify target facilities within the airport; distinguishing between different types of target facilities helps in classifying and assessing the unique damage patterns exhibited by these different types of target facilities.
[0055] This application improves the robustness and accuracy of airport damage assessment models in identifying target facilities by fusing multimodal data. Especially under extreme weather or nighttime observation conditions, it fully leverages the stability advantages of radar imagery to compensate for the shortcomings of optical imagery, thereby achieving more comprehensive feature extraction and semantic understanding.
[0056] Step S1132: Perform feature fusion segmentation processing on the deep semantic features and shallow fusion features corresponding to all target facilities to determine the segmentation mask image.
[0057] In some implementations, a cascaded decoder is used to perform feature fusion segmentation processing on the deep semantic features and shallow fusion features corresponding to all target facilities to determine a segmentation mask map. Among these, A cascaded decoder is a module based on the U-Net architecture. It gradually restores the resolution of deep semantic features through upsampling operations and then fuses and segments the upsampled data with shallow fusion features. In this application, the cascaded decoder based on the U-Net architecture can simultaneously preserve semantic information and boundary details during segmentation, thereby improving the accuracy of the final segmentation result.
[0058] Feature fusion segmentation refers to integrating features from different levels and outputting pixel-level segmentation results through a segmentation head. In feature fusion segmentation, the airport damage assessment model stitches together features at different scales to ensure the coherence and integrity of information transmission.
[0059] In some implementations, deep semantic features and shallow fusion features are input into the airport target segmentation decoder module, and the segmentation mask map is determined by the airport target segmentation decoder module.
[0060] In this embodiment, firstly, an attention mechanism is introduced based on shallow features to enhance the airport damage assessment model's focus on important features, thereby extracting more representative deep semantic features. These deep semantic features not only contain category information of the target facilities but also reflect their spatial distribution characteristics, which helps improve the accuracy of the segmentation mask. Next, an upsampling operation and multi-level feature fusion are implemented through a cascaded decoder, enabling the final output segmentation mask to have higher resolution and detail representation capabilities, significantly improving the segmentation effect of airport target facilities.
[0061] In some embodiments, step S1131 includes steps S200 to S203: Step S200: Based on the self-attention calculation performed on the first shallow features corresponding to all target facilities, determine the first attention information corresponding to all target facilities; Here, self-attention computation is a deep learning technique for processing sequence data. It dynamically adjusts the degree of attention to each element by calculating the correlation weight between each element in the sequence and all other elements, thereby capturing the complex dependencies within the sequence.
[0062] In this application, the self-attention computing mechanism is applied to the first shallow features of each target facility extracted from optical images, which can enhance the expressive power of the internal structure of the first shallow features of each target facility extracted from optical images.
[0063] In some implementations, a multi-head self-attention mechanism can be used to calculate the first shallow features corresponding to each of the target facilities to obtain the first attention information corresponding to each of the target facilities.
[0064] In implementation, the first shallow features corresponding to all target facilities are divided into multiple heads. Through a multi-head self-attention mechanism, dependencies in the sequence can be captured simultaneously from multiple perspectives. Each head may focus on capturing different types of dependencies, such as syntactic structure, semantic relationships, long-distance dependencies, etc.
[0065] Understandably, by calculating attention information for the first shallow features corresponding to all target facilities, the system can focus on key areas in the optical image (such as runway cracks and apron edges) and ignore irrelevant parts of the optical image, thereby improving the accuracy and robustness of feature representation.
[0066] Step S201: Based on the self-attention calculation performed on the second shallow features corresponding to all target facilities, determine the second attention information corresponding to all target facilities; In this application, the self-attention computing mechanism is applied to the second shallow features of each target facility extracted from radar images, which can enhance the expressive ability of the internal structure of the second shallow features of each target facility extracted from radar images and strengthen the spatial distribution and structural information of the target facilities in radar images.
[0067] In some implementations, a multi-head self-attention mechanism can be used to calculate the second shallow features corresponding to each of the target facilities to obtain the second attention information corresponding to each of the target facilities.
[0068] In implementation, the second shallow features corresponding to all target facilities are divided into multiple heads. Through a multi-head self-attention mechanism, dependencies in the sequence can be captured simultaneously from multiple perspectives. Each head may focus on capturing different types of dependencies, such as syntactic structure, semantic relationships, long-distance dependencies, etc.
[0069] It is understandable that by calculating the attention information of the second shallow features corresponding to all target facilities, the system can focus on the outline and change points of the target facilities in the radar image, improve the spatial description accuracy of the target facilities in the radar image, and thus enhance the complementarity with optical image, achieving more comprehensive feature fusion.
[0070] Step S202: Based on the cross-attention calculation between the first shallow features and the second shallow features corresponding to all target facilities, determine the third attention information corresponding to all target facilities. Here, cross attention is a deep learning mechanism that extends the idea of self attention, allowing the model to dynamically focus on information from different sequences or modalities when processing sequential or modal data.
[0071] In this application, cross-attention computation is used to connect the first shallow features of the target facility extracted from optical images and the second shallow features of the target facility extracted from radar images. This enables the system to understand the correlation between the same target facilities in different modalities. By using the cross-attention computation method, this application can better integrate information from the two modalities, discover details that may be missed in optical images (such as obscured runway sections), and at the same time, utilize the stable spatial structure in radar images to correct noise or blurred areas in optical images.
[0072] In some implementations, cross-attention computation uses a query-key-value mechanism, employing a first shallow feature of the target facility extracted from optical imagery as the query and a second shallow feature of the target facility extracted from radar imagery as the key and value, thereby obtaining the interaction information between the two. This process helps generate more semantically consistent fused features, enhancing the model's adaptability to complex scenarios.
[0073] It is understandable that there is some semantic overlap and difference between the first shallow features of target facilities extracted from optical images and the second shallow features of target facilities extracted from radar images. The cross-attention mechanism can dynamically adjust the information weights based on the similarity between the first and second shallow features, thereby establishing a semantic connection between the first and second shallow features, realizing the complementarity of multi-source data, and thus enabling higher-precision airport target segmentation and damage assessment.
[0074] Step S203: Based on the first attention information, second attention information and third attention information corresponding to all target facilities, determine the deep semantic features corresponding to all target facilities.
[0075] Here, deep semantic features refer to high-dimensional feature representations formed after multiple rounds of feature fusion and attention mechanisms. These high-dimensional feature representations contain rich semantic information and spatial contextual relationships.
[0076] It should be noted that the first attention information emphasizes the local structural features of the first shallow features of the target facility extracted from optical images, representing the single-modal features of the optical images; the second attention information highlights the spatial stability of the second shallow features of the target facility extracted from radar images, representing the single-modal features of the radar images; and the third attention information further integrates the relationship between the first and second shallow features, providing cross-modal global information. Combining the first, second, and third attention information in this way generates richer and more stable deep semantic features, improving the airport damage assessment model's adaptability and accuracy in complex environments, thereby enabling the airport damage assessment model to accurately locate and assess the damage status of airport target facilities.
[0077] In this embodiment, a self-attention mechanism is used to model the shallow features of optical and radar images separately, which enhances the semantic consistency within a modality. Simultaneously, a cross-attention mechanism establishes connections between cross-modal features, thereby achieving more comprehensive information interaction. This multi-attention mechanism design enables the model to better understand the contextual relationships of the target facility, improving its sensitivity to local details and global structure, thus generating higher-quality deep semantic features and providing strong support for subsequent segmentation tasks.
[0078] In some embodiments, step S1132 above includes steps S210 and S211: Step S210: Based on the deep semantic features corresponding to all target facilities and the shallow fusion features corresponding to all target facilities, obtain the fusion features of the airport; Deep semantic features refer to high-dimensional abstract features extracted through deep neural networks (such as ResNet), which can express the semantic information of target facilities, such as the straightness of a runway and the rectangular outline of an apron. Deep semantic features usually have stronger discriminative power, but retain less detailed structure. Shallow fusion features refer to features obtained by fusing low-level features (such as edges and textures) of optical images and SAR images across modalities through attention mechanisms or weighted summation methods, which have better local information preservation capabilities. In this application, the deep semantic features corresponding to all target facilities are combined with the shallow fusion features corresponding to all target facilities, which can simultaneously preserve semantic information and local details, thereby improving the accuracy and robustness of segmentation.
[0079] In some implementations, an upsampling operation is used to gradually restore the resolution of deep semantic features, and shallow fusion features are received through skip connections. In this way, the feature vectors corresponding to the upsampled deep semantic features and shallow fusion features are merged into a single feature vector through a concatenation operation, thereby realizing multi-scale feature fusion.
[0080] Step S211: The fusion features are segmented according to the type of target facility of the airport using the segmentation head to obtain a segmentation mask map.
[0081] Here, the segmentation head is part of the decoder module. Its main function is to output pixel-level segmentation results based on the input fusion features. In this application, segmentation is performed according to the type of the target facility, which means that different categories (such as runways, aprons, hangars, etc.) are distinguished during the segmentation process, and different label values are assigned to these categories to form a pixel-level segmentation mask map of multiple airport types.
[0082] In practical implementation, the segmentation head can use the U-Net structure or other variants, such as the Feature Pyramid Network (FPN), to enhance the model's contextual understanding and boundary detection accuracy. For example, in a remote sensing image of an airport, the segmentation head can determine whether a certain area is a runway or an apron based on fused features, and then represent these areas with different colors or values. Finally, the segmentation head outputs a segmentation mask map of the airport containing all key facilities and their corresponding category information.
[0083] In this application, firstly, deep semantic features are combined with shallow fusion features to form richer feature representations, which helps improve the airport damage assessment model's ability to perceive changes in target boundaries and structures. Then, by classifying the fusion features through a segmentation head, segmentation mask maps for different types of target facilities can be generated, thereby achieving fine-grained segmentation of key airport facilities and providing a reliable foundation for subsequent damage assessment.
[0084] In some embodiments, step S210 may include steps S2101 and S2102: Step S2101: For each target facility, the upsampling result obtained after upsampling based on the deep semantic features of the target facility is fused with the shallow features of the target facility to obtain the multi-layer fusion features of the target facility. Here, upsampling is the process of restoring low-resolution feature maps to high resolution, with the aim of making the spatial dimensions of deep semantic features consistent with the original input image, facilitating subsequent pixel-level segmentation operations. Upsampling methods include transposed convolution, interpolation (such as bilinear interpolation), and pyramid pooling.
[0085] By upsampling the deep semantic features, we can not only improve the spatial resolution of the deep semantic features, but also retain the semantic information of the deep semantic features, so that each pixel can correspond to the specific facility location and status.
[0086] Shallow fusion features are features obtained by fusing multimodal data (such as optical and SAR images) in early network layers, mainly including local information such as edges, textures, and shapes. Shallow fusion features can leverage the complementarity between different modalities to improve the perception of blurred or occluded areas.
[0087] Multi-layer fusion features combine upsampled deep semantic features with shallow fusion features, enabling cross-scale and cross-modal information interaction. While maintaining high semantic expressiveness, multi-layer fusion features preserve fine-grained spatial information, thereby improving the accuracy and robustness of target segmentation. Multi-layer fusion features are a key component in achieving refined segmentation of airport target facilities (such as runways, aprons, and oil depots).
[0088] Step S2102: Determine the airport's fusion characteristics based on the multi-layer fusion characteristics corresponding to all target facilities.
[0089] Here, the airport's fusion characteristics refer to the comprehensive characteristic representation formed by globally integrating the multi-layered fusion characteristics of all target facilities.
[0090] In some implementations, the multi-layered fusion features of each target facility can be integrated using methods such as weighted summation, attention mechanisms, or hierarchical aggregation to obtain the fusion features of the airport. The advantage of the above methods is that, on the one hand, they can preserve the feature information of each target facility, and on the other hand, they can improve the distribution of target facilities in the airport through global modeling.
[0091] In this embodiment, firstly, deep semantic features are restored to a higher resolution through upsampling, and then concatenated with shallow fusion features to form multi-layer fusion features, which can retain more spatial detail information and improve the segmentation effect. Then, by globally modeling the multi-layer fusion features of each target facility, the fusion features of the airport are further generated, providing data support for the generation of the segmentation mask map.
[0092] In some embodiments, step S130 may include steps S131 and S132: Step S131: Based on the comparison results of each target facility before and after the damage and the preset level assessment strategy, determine the damage level of each target facility. Here, damage level refers to the degree of damage categorized according to a pre-defined standard system, based on changes in the target facility's appearance in pre- and post-disaster images. For example, damage levels can be divided into four categories: no damage, minor damage, severe damage, and complete destruction. The damage assessment strategy is a set of pre-defined rules that guide how changes to the target facility are mapped to specific damage levels. The rules included in the damage assessment strategy may be based on multiple indicators such as area change rate, outline integrity, and structural damage degree, and are quantified using machine learning models.
[0093] Step S132: Determine the damage level mask of the airport based on the damage level and the segmentation mask of the airport.
[0094] Here, the damage level mask is a pixel-level output. Each pixel value in the damage level mask corresponds to the damage level of the target facility at that location. The damage level mask is generated by further overlaying damage level information onto the segmentation mask.
[0095] In some implementations, the location of each target facility within the airport is identified by segmentation masking; then, the damage level of each target facility is assigned to the corresponding module in the segmentation masking, ultimately generating a complete damage level masking of the airport, which is used for subsequent emergency response or repair planning.
[0096] In this embodiment, firstly, by comparing the changes in target facilities in pre-disaster and post-disaster optical images and combining this with a preset damage level assessment strategy, a quantitative analysis of the damage extent of each target facility can be achieved. Then, by combining the segmentation mask map and damage level information, a damage level mask map is generated, enabling a comprehensive assessment of the overall damage status of the airport and providing a scientific basis for emergency response and repair decisions.
[0097] In some embodiments, the above-described airport target damage assessment method based on multi-source image fusion may include steps S140 and S150: Step S140: Determine the similarity metric loss and focus loss between the segmentation mask image and the first reference information respectively; determine the first loss information based on the focus loss and similarity metric loss; the first loss information is used to describe the calculation accuracy of the airport damage assessment model for the location information and category information of each target facility; the first reference information is the real location information and real type information of each target facility in the airport; Here, the primary reference information is high-precision, manually labeled data, including the actual location coordinates and category labels of each target facility, serving as the standard answer for training and validating the airport damage assessment model. Similarity metric loss is a loss function that measures the similarity between the segmentation mask image output by the airport damage assessment model and the actual annotations; examples include Dice loss. Focus loss is a loss function designed to address class imbalance problems. By assigning higher weights to hard-to-classify samples, it makes the airport damage assessment model pay more attention to easily overlooked target details, improving the overall segmentation performance.
[0098] In this application, the similarity measurement loss and focus loss are combined to comprehensively evaluate the performance of the segmentation mask image in both localization and classification dimensions. The resulting first loss information can serve as an important basis for optimizing the airport damage assessment model and guide the airport damage assessment model to further improve the recognition accuracy of each target facility in subsequent iterations.
[0099] The method for determining the similarity measurement loss is shown in the following formula (1): (1); in, M Total number of pixels K Indicates the total number of categories. Represents pixels j Category k The predicted probability, Represents pixels j truth labels, L dice This represents the loss in similarity measurement.
[0100] The method for determining focus loss is shown in the following formula (2): (2); in, L focal This indicates a loss at the focal point.
[0101] The method for determining the first loss information is shown in the following formula (3): (3); in, λ focal It is control L focal Scaling factor of weights L seg This is the first piece of information indicating a loss.
[0102] Step S150: Determine the cross-entropy loss between the damage level mask and the second reference information; based on the cross-entropy loss and the first loss information, determine the second loss information; the second loss information is used to describe the calculation accuracy of the airport damage assessment model for the damage level of each target facility; the second reference information is the actual damage information of each target facility of the airport.
[0103] Here, the second reference information consists of actual damage level data annotated by experts. This second reference information is provided to the airport damage assessment model as a basis for judging the accuracy of the prediction results. Cross-entropy loss is a commonly used loss function in classification tasks, used to measure the difference between the damage level predicted by the airport damage assessment model and the actual annotation.
[0104] In this application, by combining cross-entropy loss with the first loss information, a comprehensive evaluation of the airport damage assessment model can be achieved in three dimensions: target location, category identification, and damage level classification, thereby generating more comprehensive and accurate second loss information.
[0105] The method for determining the cross-entropy loss is shown in the following formula (4): (4); in,L ce This represents the cross-entropy loss.
[0106] The method for determining the second loss information is shown in the following formula (5): (5); in, λ ce It is control L ce Scaling factor of weights L cls This is the second piece of information regarding the loss.
[0107] In this embodiment, by introducing multiple loss functions, the airport damage assessment model is comprehensively optimized in target segmentation and damage level assessment tasks. This enables refined training and optimization of the airport damage assessment model in multiple dimensions, thereby improving the robustness and generalization ability of the airport damage assessment model in complex scenarios, and thus significantly enhancing the overall performance and reliability of airport target damage assessment.
[0108] The following describes the application of the embodiments of this application in a real-world scenario.
[0109] Airports are crucial transportation hubs and important nodes in public safety, critical infrastructure protection, disaster emergency response, and airport operation management, playing an irreplaceable role in both civilian and specific application areas. However, airports are highly vulnerable to damage during extreme events, sabotage, or natural disasters, and their damage assessment directly impacts emergency response efficiency and post-disaster recovery. Traditional airport damage assessment primarily relies on manual on-site surveys or single-modal remote sensing image analysis, which suffers from low efficiency, high risk, and insufficient accuracy. For example, manual surveys are time-consuming and labor-intensive, and difficult to conduct in high-risk or inaccessible environments; single-modal data is limited by sensor characteristics, and relying solely on single-modal data cannot comprehensively utilize the complementarity of multi-source data, making it difficult to fully capture the damage characteristics of airport targets. In the field of airport damage assessment, existing assessment models are very limited and mostly focus on macroscopic analysis of the overall damage level of the airport, lacking fine-grained segmentation and independent assessment of key internal airport facilities (such as runways, aprons, hangars, aircraft sheds, and fuel depots). Traditional methods typically utilize single-modal remote sensing imagery and employ change detection or classification models (such as support vector machines and random forests) to determine whether an airport is damaged overall. However, such methods have significant limitations: (1) Coarse target segmentation: Existing studies are mostly unable to accurately distinguish different functional areas inside the airport, resulting in deviations in damage location.
[0110] (2) Simplified damage classification: Most methods only detect the changed areas of the airport, ignoring the quantitative classification of the degree of damage, which seriously affects the repair decision.
[0111] (3) Lack of multi-objective collaborative analysis: As a complex system, the function of an airport depends on the coordinated operation of multiple sub-objectives. Once a key sub-objective in an airport is damaged, it will seriously affect the airport's support function. However, existing methods have not established a sub-objective evaluation function, which affects the reliability of the evaluation results and causes the evaluation results to deviate from the actual application scenario.
[0112] With the development of multimodal deep learning technology and the widespread application of multi-source remote sensing data (such as optical imagery, radar imagery, infrared imagery, etc.), this type of method has partially alleviated the limitations of single-modal methods by fusing multi-source data. However, it still has the following problems: (1) High computational cost: The model has a large number of parameters, high training and inference resource consumption, and is difficult to deploy on edge devices or in wartime environments.
[0113] (2) Lack of fine-grained information: Insufficient alignment of multimodal features leads to inadequate extraction of local damage features (such as runway cracks and building structure collapse).
[0114] (3) Low positioning accuracy: Location prediction relies on coarse-grained features and cannot accurately measure the spatial distribution of the damaged area.
[0115] Furthermore, the current field of airport damage assessment faces significant data bottlenecks. Existing publicly available datasets have obvious limitations in terms of target coverage and annotation granularity. Taking the internationally recognized xView2 Building Damage Assessment (xBD) dataset as an example, although it constructs a building damage detection benchmark using multi-temporal satellite imagery and provides multi-dimensional annotations including damage severity grading (undamaged, slightly damaged, severely damaged, completely damaged), its research scope is limited to civil building targets, making it difficult to support the refined assessment needs of complex facilities like airports. This data gap is particularly pronounced in the remote sensing field under extreme events. Currently, there is no publicly available remote sensing dataset specifically for airport target damage assessment benchmarks, severely limiting the development and validation of related algorithms. It is worth noting that even in non-damage assessment fields, existing airport remote sensing datasets (such as WHU-RS19) only contain coarse-grained classification annotations, lacking pixel-level semantic segmentation annotations for airport facility components. Therefore, constructing a specialized airport target damage assessment dataset has become a key infrastructure project to overcome algorithm development bottlenecks and realize the evolution of airport damage assessment towards "quantification, intelligence, and real-time" capabilities.
[0116] To address the aforementioned issues, this application proposes a multi-source image fusion-based method for airport target damage assessment. By deeply fusing multi-source data, multi-modal data interaction is achieved, comprehensively extracting damage features. Simultaneously, the model enhances cross-temporal and cross-modal local feature alignment, accurately capturing detailed information such as runway uplift and building collapse. Furthermore, by optimizing the model structure, this method can better analyze the location and damage level of damaged areas, providing an efficient, accurate, and adaptable solution for airport target damage assessment. It can be widely applied in damage effect assessment, disaster emergency response, and infrastructure maintenance, enhancing adaptability to complex scenarios.
[0117] The airport target damage assessment method based on multi-source image fusion proposed in this application may include steps S1 to S4: Step S1: Data preparation and annotation; In some implementations, data preparation and annotation may include steps S11 and S12: Step S11: Acquire airport data from remote sensing images; This application constructs a specialized dataset for airport damage assessment. This specialized dataset may include high-resolution optical imagery (i.e., the first optical imagery mentioned above) and synthetic aperture radar (SAR) imagery (i.e., the radar imagery mentioned above). These specialized datasets may be airports located in various locations, and the acquisition time of the imagery data for these airports may differ.
[0118] Step S12: Airport data annotation.
[0119] First, remote sensing experts with experience in airport target identification developed an assessment system for identifying typical airport targets and assessing hard damage. Simultaneously, a professional dataset was used to label the location and damage status of multiple facilities (such as runways, aprons, and hangars) at different airports. The damage status was categorized into four levels: no damage, minor damage, severe damage, and complete destruction. It should be noted that this professional dataset can include open-source imagery data, such as OpenStreetMap (OSM) data, or licensed imagery data, such as airport imagery from Google Maps.
[0120] Step S2: Data augmentation; To improve the model's generalization ability in airport target damage assessment tasks, sampling data augmentation methods are used to preprocess the data. These data augmentation methods may include, but are not limited to, the following: Method 1: Geometric transformations, such as horizontal flipping, vertical flipping, rotation, cropping, and multi-image stitching, optimize the visual information of the image. Each geometric transformation is applied simultaneously to the image and the corresponding target segmentation mask to ensure that the spatial relationship between the target facility and the airport background remains unchanged.
[0121] Method 2: Color enhancement, for example, by adjusting the brightness, contrast, saturation and sharpness of the image, the changes in the image under different lighting and weather conditions are simulated, thereby enhancing the model's adaptability to environmental changes.
[0122] Step S3: Construction and training of airport target damage assessment model based on multi-source image fusion.
[0123] like Figure 2 As shown, the airport target damage assessment model (i.e., the aforementioned airport damage assessment model) includes a shallow feature extraction and cross-modal fusion module 1, a deep feature extraction and cross-modal fusion module 2, an airport target segmentation decoder module 3, and a cross-temporal feature fusion and damage level assessment module 4. Among these, 1) The shallow feature extraction and cross-modal fusion module includes a dual-branch convolutional neural network (CNN) architecture and a shallow feature fusion structure. Among them, A dual-branch convolutional neural network architecture is used to extract the first shallow features of optical images and the second shallow features of SAR images. The backbone network of the dual-branch convolutional neural network architecture uses ResNet (residual neural network) to capture edge, texture and shape information (such as runway straight lines and apron rectangular outlines) through multi-scale convolutional layers.
[0124] Shallow feature fusion structures are used to generate shallow fused features by performing feature fusion processing (e.g., feature weighting) on the output of a bi-branch convolutional neural network architecture. These structures can include global average pooling and attention-based units, such as different versions of the Squeeze-and-Excitation (SE) module. It is understood that shallow feature fusion structures can achieve data complementarity between optical and radar images; for example, in cloud-covered scenarios, they can effectively utilize the penetrating power of SAR to supplement features in areas missing from optical data.
[0125] 2) The deep feature extraction and cross-modal fusion module includes units built based on an attention mechanism and a global semantic enhancement layer. Among them, The unit built based on the attention mechanism can include self-attention mechanism, cross-attention mechanism, etc. The input of the unit is the first shallow feature of the optical image and the second shallow feature of the SAR image. In this way, on the one hand, the self-attention mechanism in the unit performs multi-head self-attention calculation on the first shallow feature and the second shallow feature respectively to enhance intramodal semantic consistency; on the other hand, the interaction relationship between different modal features (first shallow feature and second shallow feature) is calculated by using the cross-attention mechanism in the unit.
[0126] The global semantic enhancement layer is used to perform secondary enhancement and optimization on the output of units built based on the attention mechanism, generate deep semantic features, and improve robustness in complex environments.
[0127] 3) Airport target segmentation decoder module. This module adopts a cascaded decoder architecture, with inputs being the obtained deep semantic features and shallow fused features. Among them, The decoder is based on the U-Net structure. It gradually restores the resolution of deep semantic features through upsampling and skips connections with the shallow feature extraction and cross-modal fusion modules to achieve multi-layer fusion decoding of shallow fusion features and deep semantic features (e.g., simple splicing and superposition fusion). Finally, it outputs an airport target segmentation mask map (i.e. the segmentation mask map mentioned above) through the segmentation head. This airport target segmentation mask map is used to characterize the segmentation of different facilities in the airport.
[0128] 4) Cross-temporal feature fusion and damage level assessment module: The input is high-resolution airport images before and after the disaster. Since the target outline is blurred in the post-disaster images, in order to refine the damage assessment results, the airport target location information (i.e., airport target segmentation mask map) output in the first stage is used as an input for the damage level assessment stage to obtain a more accurate damage level mask map.
[0129] The cross-temporal feature fusion and damage level assessment module includes: a dual-branch hierarchical coding structure based on Transformer, a cross-temporal feature fusion unit based on attention mechanism and convolutional blocks, and a multi-scale decoder based on CNN. The output of the cross-temporal feature fusion and damage level assessment module corresponds to the airport target damage assessment results. The cross-temporal feature fusion unit combines deep and shallow features to enhance information transmission and fusion, while also achieving weighted fusion between features of different scales to enhance the interaction of multi-scale information.
[0130] In some implementations, to optimize the model's performance in multi-source image fusion and damage assessment tasks, end-to-end training can be achieved using a constructed dataset through gradient backpropagation and parameter updates. Specifically, target segmentation loss and damage classification loss include cross-entropy loss, Dice loss, and focus loss. Loss calculations optimize target boundary accuracy and enhance detail capture capabilities, while mitigating class imbalance and reducing intermodal feature differences. The total loss function uses a dynamic weight allocation strategy to balance the contributions of each task, ensuring collaborative optimization of the model in multi-task learning.
[0131] Step S4: Reasoning and testing of airport target damage assessment model based on multi-source image fusion.
[0132] In some implementations, the reasoning and testing process of the airport target damage assessment model based on multi-source image fusion may include steps S41 to S45: Step S41: Input unlabeled airport images of different modalities into the trained airport target damage assessment model.
[0133] Step S42: The airport target damage assessment model first extracts local features (such as edges and textures) from optical and SAR images through a shallow feature extraction module (i.e., the aforementioned dual-branch convolutional neural network architecture), and then achieves cross-modal information complementarity through a shallow feature fusion module (i.e., the aforementioned shallow feature fusion architecture) to obtain shallow fused features. Finally, a deep feature extraction and fusion module generates a fused high-dimensional semantic feature representation (i.e., the aforementioned deep semantic feature).
[0134] Step S43: The airport target damage assessment model inputs the deep semantic features into the cascaded decoder (i.e., the airport target segmentation decoder module mentioned above), and concatenates them with the shallow fusion features through upsampling operations to restore the high-resolution segmentation result and output the airport target segmentation mask map.
[0135] Step S44: The second stage inputs pre-disaster and post-disaster images of the same modality (i.e., multiple first optical images and multiple second optical images mentioned above), and simultaneously uses the segmentation result of the airport target location obtained in the first stage (i.e., the airport segmentation mask map mentioned above) as an input for the damage level assessment stage. The cross-temporal feature fusion unit combines features from the pre-disaster and post-disaster images to establish the connection between shallow and deep features in the dual-branch network. The damage type classifier identifies the degree of damage to airport facilities and finally outputs a mask map representing the damage level.
[0136] Step S45: Performance Evaluation and Testing. The model performance can be evaluated using a test set derived from the constructed dataset. For example, metrics such as mean intersection-over-union ratio (NICU) and F1 score can be used to measure the accuracy of airport target segmentation. Harmonic mean (HRM) is used to evaluate classification results, mitigating class imbalance and assessing the model's performance in target damage classification. The evaluation process aims to verify the model's accuracy, reliability, and robustness in real-world applications, ensuring its practicality under complex environmental requirements.
[0137] In some embodiments, such as Figure 3 The training process of the airport target damage assessment model is illustrated, and may include steps S300 to S330: Step S300: Acquire multimodal airport images at different time points before and after the disaster, and label key facility targets in the multimodal airport images, including target category, facility location, and target damage level.
[0138] Step S310: Use data augmentation methods such as geometric transformation and color enhancement, randomly select different combinations of enhancement strategies, and augment the obtained labeled data.
[0139] Step S320: Use the constructed multimodal airport dataset to train the airport target damage assessment model based on multi-source image fusion; Step S330: Input the unlabeled multimodal airport image into the trained model to obtain a mask map representing the damage level of the airport target (i.e., the damage level mask map mentioned above).
[0140] In some embodiments, such as Figure 4 The application process of the airport target damage assessment model is illustrated, and may include step S400: Step S400: Input the multimodal, multi-temporal airport images into the airport target damage assessment model based on multi-source image fusion, and output a multi-class mask map (i.e., the damage level mask map mentioned above) representing the damage level of different target facilities.
[0141] This application proposes an airport target damage assessment method based on multi-source image fusion, which can bring the following beneficial effects: 1. Task Decoupling and Accuracy Improvement: Traditional methods typically couple target localization and damage assessment within the same network, leading to mutual interference between tasks and affecting assessment accuracy. This application implements airport facility target localization and damage assessment classification separately through a two-stage network architecture. Combining pre-disaster and post-disaster multimodal data, a cross-stage feature reuse strategy is adopted. The high-precision segmentation results generated in the first stage guide the refinement of damage areas in the second stage, improving the accuracy and robustness of target segmentation and damage assessment.
[0142] 2. Optimization of Multimodal Data Fusion Capabilities: Traditional methods typically employ simple weighted summation or splicing strategies when fusing multimodal data, making it difficult to fully exploit the complementary information between modalities. In the first stage, this application utilizes pre-disaster multimodal data for shallow and deep feature extraction and cross-modal fusion, capturing fine-grained information and fully exploiting the complementarity of cross-modal features. This enables multi-level fusion of optical imagery and SAR data, improving the model's adaptability to complex scenes.
[0143] 3. By adopting skip connections and multi-layer cascaded decoders, multi-layer fusion decoding of shallow and deep features is achieved, which improves segmentation accuracy and decoding capability, and generates high-precision airport target segmentation results.
[0144] 4. Enhanced Cross-Temporal Feature Interaction Capability: Existing methods often rely on single-temporal data or simple stitching of pre- and post-war images, lacking effective utilization of cross-temporal features. This application introduces a cross-temporal feature fusion and decoding module to combine deep and shallow features from images of different temporal phases, enhancing information transmission and fusion. Simultaneously, it achieves weighted fusion between features at different scales, not only enhancing the interaction of multi-scale information but also increasing sensitivity to local damage and global structural changes. This addresses the problem of insufficient utilization of cross-temporal features in traditional methods.
[0145] 5. Construction of a Refined Dataset: Existing publicly available datasets have significant limitations in terms of target coverage and annotation precision, especially in airport target damage detection tasks, lacking refined annotations for critical facilities such as runways, aprons, and hangars. This application constructs a dedicated benchmark dataset for airport target damage detection based on multi-temporal and multi-modal satellite imagery. This dataset provides a multi-dimensional annotation system covering target categories (such as runways, hangars, and fuel depots), target locations, and damage severity classifications (mild, moderate, and severe), filling the gap in high-quality annotation data in this field and promoting the development of related technologies towards refinement and practical application.
[0146] 6. A joint optimization loss function was designed, which combines segmentation loss and classification loss. The model parameters are optimized through multi-task learning, which further improves the model's positioning accuracy and stability, and alleviates the impact of uneven damage level distribution on model performance.
[0147] Based on the foregoing embodiments, this application provides an airport target damage assessment device based on multi-source image fusion, such as... Figure 5 As shown, the device 500 includes: The acquisition module 501 is used to acquire, in response to damage to the airport, multiple first optical images and multiple radar images of the airport before the damage, and multiple second optical images of the airport after the damage. The first processing module 502 is used to obtain a segmentation mask map of the airport by performing feature extraction and fusion processing on multiple first optical images and multiple radar images based on a pre-built airport damage assessment model; the segmentation mask map is used to characterize the location information and type information of each target facility in the airport; The second processing module 503 is used to obtain the comparison results of each target facility before and after damage by comparing each target facility in multiple first optical images and multiple second optical images based on the airport damage assessment model. The third processing module 504 is used to perform damage level assessment processing based on the airport damage assessment model by comparing the results of each target facility before and after damage and the segmentation mask map of the airport, and obtain the damage level mask map of the airport; the damage level mask map is used to characterize the degree of damage of each target facility.
[0148] In some embodiments, the first processing module includes: The first determining unit is used to perform feature extraction processing on multiple first optical images and multiple radar images using a dual-branch convolutional neural network architecture to determine the first shallow features of each target facility in the multiple first optical images and the second shallow features of each target facility in the multiple radar images; the first shallow features or the second shallow features are used to characterize the attribute information of each target facility. The second determining unit is used to perform feature fusion processing on the first shallow feature and the second shallow feature to determine the shallow fusion features of each target facility. The third determining unit is used to determine the segmentation mask image based on the first shallow feature, the second shallow feature, and the shallow fusion feature corresponding to all target facilities in the airport.
[0149] In some embodiments, the third determining unit includes: The first determining subunit is used to determine the deep semantic features corresponding to all target facilities based on attention calculations of the first shallow features and the second shallow features corresponding to all target facilities respectively. The second determining subunit is used to perform feature fusion and segmentation processing on the deep semantic features and shallow fusion features corresponding to all target facilities respectively, and to determine the segmentation mask map.
[0150] In some embodiments, the first determining subunit includes: determining first attention information corresponding to each of the target facilities based on self-attention calculation performed on first shallow features corresponding to each of the target facilities; determining second attention information corresponding to each of the target facilities based on self-attention calculation performed on second shallow features corresponding to each of the target facilities; determining third attention information corresponding to each of the target facilities based on cross-attention calculation performed between the first shallow features and the second shallow features corresponding to each of the target facilities; and determining deep semantic features corresponding to each of the target facilities based on the first attention information, second attention information, and third attention information corresponding to each of the target facilities.
[0151] In some embodiments, the second determining subunit includes: obtaining the airport's fusion features based on the deep semantic features corresponding to all target facilities and the shallow fusion features corresponding to all target facilities; and segmenting the fusion features according to the type of the airport's target facilities using a segmentation head to obtain a segmentation mask map.
[0152] In some embodiments, the second determining subunit further includes: for each target facility, obtaining a multi-layer fusion feature of the target facility by combining the upsampling result obtained after upsampling processing based on the deep semantic features of the target facility with the shallow fusion features of the target facility; and determining the fusion feature of the airport based on the multi-layer fusion features corresponding to all target facilities respectively.
[0153] In some embodiments, the third processing module includes: The fourth determination unit is used to determine the damage level of each target facility based on the comparison results of each target facility before and after the damage and the preset level assessment strategy. The fifth determining unit is used to determine the damage level mask map of the airport based on the damage level and the segmentation mask map of the airport.
[0154] In some embodiments, the apparatus further includes: The first determining module is used to determine the similarity metric loss and focus loss between the segmentation mask image and the first reference information, respectively; and to determine the first loss information based on the focus loss and similarity metric loss; the first loss information is used to describe the calculation accuracy of the airport damage assessment model for the location information and category information of each target facility; the first reference information is the real location information and real type information of each target facility in the airport; The second determination module is used to determine the cross-entropy loss between the damage level mask map and the second reference information; based on the cross-entropy loss and the first loss information, the second loss information is determined; the second loss information is used to describe the calculation accuracy of the airport damage assessment model for the damage level of each target facility; the second reference information is the actual damage information of each target facility of the airport.
[0155] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These software products are stored in a storage medium and include several instructions to cause a motor vehicle to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination. This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement any of the methods described above. This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. The computer-readable storage medium can be transient or non-transient. This application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement some or all of the steps in any of the above-described methods. The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc. It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding. The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for airport target damage assessment based on multi-source image fusion, characterized in that, The method includes: In response to damage to the airport, multiple first optical images and multiple radar images of the airport before the damage, and multiple second optical images of the airport after the damage; Based on a pre-built airport damage assessment model, a segmentation mask map of the airport is obtained by extracting and fusing features from the multiple first optical images and the multiple radar images; the segmentation mask map is used to characterize the location and type information of each target facility in the airport. Based on the airport damage assessment model, by comparing each target facility in the plurality of first optical images and the plurality of second optical images, the comparison results of each target facility before and after damage are obtained. Based on the airport damage assessment model, a damage level mask map of the airport is obtained by comparing the results of each target facility before and after damage and by performing damage level assessment on the segmentation mask map of the airport; the damage level mask map is used to characterize the degree of damage to each target facility.
2. The method based on claim 1, characterized in that, The step of obtaining a segmentation mask map of the airport by performing feature extraction and fusion processing on the plurality of first optical images and the plurality of radar images includes: A dual-branch convolutional neural network architecture is used to extract features of each target facility from the plurality of first optical images and the plurality of radar images to determine the first shallow features of each target facility in the plurality of first optical images and the second shallow features of each target facility in the plurality of radar images; the first shallow features or the second shallow features are used to characterize the attribute information of each target facility; The first shallow feature and the second shallow feature are subjected to feature fusion processing to determine the shallow fusion features of each target facility; The segmentation mask image is determined based on the first shallow feature, the second shallow feature, and the shallow fusion feature corresponding to all target facilities in the airport.
3. The method based on claim 2, characterized in that, The step of determining the segmentation mask image based on the first shallow feature, the second shallow feature, and the shallow fusion feature corresponding to all target facilities in the airport includes: Based on attention calculations on the first and second shallow features corresponding to all the target facilities respectively, the deep semantic features corresponding to all the target facilities are determined. The deep semantic features and shallow fusion features corresponding to all the target facilities are subjected to feature fusion segmentation processing to determine the segmentation mask image.
4. The method based on claim 3, characterized in that, The process of determining the deep semantic features corresponding to each of the target facilities based on attention calculations of the first and second shallow features corresponding to each of the target facilities includes: Based on the self-attention calculation performed on the first shallow features corresponding to all the target facilities, the first attention information corresponding to all the target facilities is determined. Based on the self-attention calculation performed on the second shallow features corresponding to all the target facilities, the second attention information corresponding to all the target facilities is determined. Based on the cross-attention calculation between the first shallow features and the second shallow features corresponding to all the target facilities, the third attention information corresponding to all the target facilities is determined. Based on the first attention information, the second attention information, and the third attention information corresponding to all the target facilities, the deep semantic features corresponding to all the target facilities are determined.
5. The method based on claim 3, characterized in that, The step of performing feature fusion segmentation processing on the deep semantic features and shallow fusion features corresponding to all the target facilities respectively, and determining the segmentation mask image, includes: Based on the deep semantic features corresponding to all the target facilities and the shallow fusion features corresponding to all the target facilities, the fusion features of the airport are obtained; The fused features are segmented according to the type of target facility of the airport using a segmentation head to obtain the segmentation mask image.
6. The method based on claim 5, characterized in that, The fusion features of the airport are obtained based on the deep semantic features and shallow fusion features corresponding to all the target facilities, respectively, including: For each target facility, the upsampling result obtained by upsampling the deep semantic features of the target facility is combined with the shallow fusion features of the target facility to obtain the multi-layer fusion features of the target facility. Based on the multi-layer fusion characteristics corresponding to all the target facilities, the fusion characteristics of the airport are determined.
7. The method according to any one of claims 1 to 6, characterized in that, The damage level assessment process, which involves comparing the damage results of each target facility before and after damage with the segmentation mask map of the airport, yields a damage level mask map of the airport. This process includes: Based on the comparison results of each target facility before and after the damage and the preset level assessment strategy, the damage level of each target facility is determined. Based on the damage level and the segmentation mask map of the airport, a damage level mask map of the airport is determined.
8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The similarity metric loss and focus loss between the segmentation mask image and the first reference information are determined respectively; based on the focus loss and the similarity metric loss, the first loss information is determined; the first loss information is used to describe the calculation accuracy of the airport damage assessment model for the location information and category information of each target facility; the first reference information is the true location information and true type information of each target facility in the airport; The cross-entropy loss between the damage level mask and the second reference information is determined; based on the cross-entropy loss and the first loss information, the second loss information is determined; the second loss information is used to describe the calculation accuracy of the airport damage assessment model for the damage level of each target facility; the second reference information is the actual damage information of each target facility of the airport.
9. An airport target damage assessment device based on multi-source image fusion, characterized in that, The device includes: The acquisition module is used to acquire, in response to damage to the airport, multiple first optical images and multiple radar images of the airport before the damage, and multiple second optical images of the airport after the damage. The first processing module is used to obtain a segmentation mask map of the airport by performing feature extraction and fusion processing on the multiple first optical images and the multiple radar images based on a pre-built airport damage assessment model; the segmentation mask map is used to characterize the location information and type information of each target facility in the airport; The second processing module is used to obtain the comparison results of each target facility before and after damage by comparing each target facility in the plurality of first optical images and the plurality of second optical images based on the airport damage assessment model. The third processing module is used to perform damage level assessment processing on the airport damage assessment model by comparing the results of each target facility before and after damage and the segmentation mask map of the airport, and obtain the damage level mask map of the airport; the damage level mask map is used to characterize the degree of damage of each target facility.
10. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 8.