Remote sensing damaged building target detection algorithm based on YOLOv8
By improving the feature extraction and loss function of the YOLOv8 model, the efficient identification of damaged buildings in remote sensing images is solved, and the accurate detection and rapid response of damaged buildings in remote sensing images is achieved, which improves the efficiency and accuracy of post-disaster rescue.
Patent Information
- Application Number
- CN202510348679.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to efficiently and accurately identify and locate damaged areas of building structures caused by disasters in remote sensing images. The traditional methods have poor safety and low efficiency, making it difficult to meet the timeliness of post-disaster rescue.
The improved YOLOv8 model is adopted, and the feature extraction capability is enhanced by introducing the CDFR module and EDCA module, combined with the T-Focal loss compound loss function to optimize object detection, and the convolution module and loss function of YOLOv8 are improved to improve the detection accuracy and efficiency of the model in complex environments.
It significantly improves the detection accuracy and efficiency of damaged buildings in remote sensing images, can quickly generate spatial distribution heat maps in disaster areas, optimize rescue resource allocation and reconstruction planning, and improve the scientific level of disaster emergency management.
Smart Images

Figure CN120279414A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a remote sensing damaged building target detection algorithm based on YOLOv8. Background Art
[0002] The detection of damaged building targets is a classic problem in the field of automatic extraction of remote sensing image information. Its core objective is to accurately identify and locate the areas of buildings damaged due to disasters in the images. This technology provides key spatial data support for accurate disaster assessment and post-disaster rescue decision-making by distinguishing the damaged features of buildings from normal ground object information.
[0003] Traditional methods rely on manual visual inspection and on-site investigation, which have defects such as poor safety and low efficiency, and are difficult to meet the timeliness requirements of post-disaster rescue. The damaged building target recognition algorithm based on remote sensing technology realizes rapid disaster perception through deep learning methods, providing key technical support for rescue resource allocation and reconstruction planning. Deep learning algorithms with convolutional neural networks (CNNs) as the core analyze complex remote sensing data through multi-layer feature abstraction mechanisms, significantly improving the detection efficiency and accuracy. Compared with traditional means, this method can quickly generate a heat map of the spatial distribution of the affected area, optimize the deployment path of rescue forces, support post-disaster reconstruction and urban planning, and effectively improve the scientific level of the disaster emergency management system. Summary of the Invention
[0004] The present invention provides a remote sensing image damaged building target detection algorithm based on YOLOv8, which has more accurate detection results in dealing with the problems of complex backgrounds of damaged buildings in remote sensing images, different scales caused by different degrees of building damage, and difficult detection of small damaged buildings.
[0005] Aiming at the deficiencies of the prior art, the present application proposes a method for detecting damaged targets in remote sensing images based on YOLOv8, including obtaining a damaged building data set, cleaning the original data to remove blurred and duplicate data, and finally annotating the obtained image data. The preprocessed damaged building image data set is divided into a training set, a validation set, and a test set, where the damaged building image data set includes damaged building images under different environmental backgrounds.
[0006] Replace the standard convolutional module Conv of YOLOv8 with the CDFR module. By adding multiple layers of convolution, the depth of the model's feature extraction is increased, effectively enhancing the model's ability to identify small features in complex visual environments. Add the EDCA module to the neck of YOLOv8. By reallocating channel weights through a mechanism similar to the "gating" mechanism, channels carrying richer discriminative information are given higher emphasis, effectively alleviating the problem of feature attenuation in deep neural networks. Use the new loss function T-Focalloss to enhance the model's detection ability for small targets and reduce the misdetection phenomenon caused by the imbalance between positive and negative samples.
[0007] Use the improved method described above to obtain a new network model, and use the training set to iteratively train the new network model to obtain the target weights. Use the test set to test the target weights to obtain the detection results of damaged building targets.
[0008] The CDFR module adopts a four-branch heterogeneous convolution architecture. The front-stage branch deploys a parallel structure of 3×3 standard convolution and 1×1 point convolution to achieve basic feature extraction. The rear-stage branch integrates a dual-path deformable convolution module to complete dynamic receptive field optimization, and combines filter response normalization (FRN) and adaptive rectified linear unit (ARelu) to form a non-linear activation system. This design realizes multi-level feature refinement through the complementary characteristics of heterogeneous convolution kernels, enhances the ability to capture geometric deformation features by using the sampling point adaptive mechanism of deformable convolution, and at the same time suppresses the abnormal propagation of gradients through the joint operation of FRN-ARelu, effectively alleviating the problem of information dissipation of small target features in deep neural networks, and comprehensively improving the feature representation efficiency of the model through the cross-level feature complementary mechanism.
[0009] The EDCA module adopts the dilated convolution technology to expand the receptive field without increasing the parameter scale or sacrificing the resolution, and efficiently captures large-scale context information. The module integrates a gating weighting mechanism to dynamically reconstruct the channel weight distribution through cross-dimensional attention coordination, and strengthens the activation response of highly discriminative feature channels. This module is deployed at the front end of the feature pyramid upsampling stage (Upsample Layer) and the back end of the detection head feature fusion layer (HeadConcat). Through a dual-path feature enhancement strategy: retaining the integrity of small target spatial details before upsampling and optimizing the cross-scale semantic combination efficiency after feature fusion, effectively improving the feature saliency of occluded areas. This structural design significantly improves the localization accuracy of the model for small targets and the representation robustness of local damaged areas.
[0010] The new loss function T-Focal loss is improved based on the Focal loss function, Tversky Loss function, NWD loss function, and IOU loss function of YOLOv8. The specific calculation formula is as follows
[0011] A multi-modal loss function framework is proposed to optimize object detection performance by integrating geometric alignment and distribution matching mechanisms. The CloU loss function is used to construct object geometric constraints, which realizes the centroid alignment and aspect ratio consistency optimization of the predicted box and the ground truth box by introducing a center point distance penalty term and an aspect ratio similarity metric. The normalized Wasserstein distance (NWD) is combined to construct a small object distribution similarity metric, which effectively alleviates the position sensitivity problem of small objects by using the spatial relationship of bounding boxes. To address the gradient drowning effect caused by the imbalance between positive and negative samples during training, a dynamic balance mechanism of Tversky Loss and Focal Loss is innovatively introduced: Tversky Loss strengthens the penalty for false negative samples through adjustable parameters, and Focal Loss suppresses the contribution of easy-to-classify samples through a modulation factor. The two cooperate to form a double attention mechanism for low-representative damaged regions. This composite loss system achieves multi-objective optimization through adaptive weight allocation, significantly improving the feature discrimination and cross-scale generalization capabilities of the model in complex scenarios.
[0012] In summary, the advantages of the present invention are as follows: A remote sensing damaged building target detection algorithm based on YOLOv8 is provided. The YOLOv8 model is improved, and the CDFR module is introduced in the feature extraction stage to construct a four-branch heterogeneous convolution architecture: in the front stage, a 3×3 standard convolution and a 1×1 point convolution are used in parallel to extract multi-scale basic features, and in the latter stage, a dual-path deformable convolution module is deployed. The capture ability of geometric deformation features is optimized through the dynamic sampling point offset mechanism, and the FRN-ARelu activation function is combined to suppress the abnormal propagation of gradients, effectively alleviating the problem of information attenuation of small target features in the deep network. At the same time, a dual-channel attention module (EDCA) is embedded in the neck of the feature pyramid. Dilated convolution is used to expand the receptive field to capture global context information, and the channel-space dual-attention gating is used to dynamically enhance the feature response of the damaged area. This module is deployed before the upsampling operation and after the feature fusion layer to achieve the collaborative optimization of small target detail retention and cross-scale semantic enhancement. Aiming at the multi-objective optimization problem in the model training process, a T-Focal loss composite loss function system is designed, which integrates the geometric constraint of the CloU loss, the distribution similarity measure of the NWD loss, and the sample balance mechanism of the Tversky-Focal loss: CloU jointly optimizes the bounding box alignment accuracy through the center distance and aspect ratio, NWD models the robustness of the small target position distribution using the Wasserstein distance, the Tversky loss strengthens the penalty for false negative samples through an adjustable coefficient, and the Focal loss suppresses the dominant effect of easy-to-classify samples through a modulation factor. The four form a multi-objective optimization system through dynamic weight allocation. Through the synergistic effect of heterogeneous feature extraction, cross-scale attention enhancement, and composite loss constraint, the feature discrimination ability and morphological preservation ability of damaged targets in complex backgrounds are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flowchart of a remote sensing damaged building target detection algorithm based on YOLOv8 provided by the present invention.
[0014] Figure 2 It is a schematic diagram of a network structure provided by the present invention.
[0015] Figure 3 It is a schematic diagram of the structure of a CDFR module provided by the present invention.
[0016] Figure 4 It is a schematic diagram of the structure of a dual-channel attention module EDCA provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings and implementation schemes in the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0019] Dataset: The dataset is a large image dataset for object detection in remote sensing images, which can be used to discover and evaluate objects in aerial images. The dataset has more than 5,000 pictures, more than 15,000 instances, and a total of 2 categories.
[0020] Use Labellmg to label the collected images with the category and location information of the objects.
[0021] Perform operations such as denoising and enhancing contrast on the collected images to improve the image quality; enhance the images by means of random cropping, rotation, scaling, etc. to increase the diversity of training samples; divide the dataset into a training set, a validation set, and a test set for model training, tuning, and evaluation. The training set, the validation set, and the test set are all composed of remote sensing images under different disaster conditions.
[0022] In this example, the lightweight version YOLOv8s of YOLOv8 is selected as the baseline model, and its model structure is improved to enhance the detection ability of damaged buildings.
[0023] In view of the deficiency of the original convolutional block of YOLOv8 in small target feature extraction, the present invention proposes a multi-level enhancement scheme: deep feature mining is realized by stacking multiple convolutional structures, local detail features are captured by using a 3×3 convolutional kernel with a stride of 2 for downsampling operation, and channel dimension compression and feature weight adaptive adjustment are combined with 1×1 convolution. The dual-stage deformable convolution module is innovatively introduced, and the sampling strategies of shallow detail and deep semantic features are respectively optimized through the dynamic offset learning mechanism, focusing on enhancing the feature representation ability of unconventional morphological targets. Finally, the channel dimension concatenation (Concat) technology is used to fuse the feature maps of five levels, and under the constraint of strictly maintaining the output dimension of H / 2*W / 2*C, the feature discrimination of dense small targets in complex backgrounds is effectively improved.
[0024] Deploy the EDCA module before the upsampling operation and after the feature fusion layer. The EDCA module adopts an innovative process to optimize features through a multi-layer information interaction mechanism. The specific workflow is as follows: Deploy the module before the feature map upsampling and after the fusion layer. First, use the normalized exponential function SoftMax to perform probabilistic weight assignment on the output result channel features to construct a channel-level attention heat map. Achieve differential enhancement and suppression of feature channels through element-wise multiplication operations, and complete residual feature fusion in combination with cross-layer identity mapping to form a feature enhancement layer with deep semantic information. The module integrates the Channel Shuffle mechanism, breaks the information barrier between channels through a random grouping-recombination strategy, and promotes cross-group feature knowledge transfer. Finally, construct a dual-path attention architecture, tensor splice the local detail features guided by spatial attention and the global semantic features enhanced by channel attention to form a multi-scale feature representation with high discriminability. This design significantly improves the model's ability to capture weakly significant targets through dynamic calibration of channel weights and cross-level feature collaboration, while suppressing the interference noise of complex backgrounds, achieving the collaborative optimization of detection accuracy and computational efficiency. The EDCA module's features in the attention feature extraction stage is the input used to calculate the output of the dual-channel attention module using Equation The spatial dimensions are represented by H and W below, while C refers to the channel dimension. Additionally denotes element-wise addition, represents the output of the ShuffleAttention segment, and represents the output features of the upper branch. Initially, process the data through the input features and then divide it into upper and random attention branches. The specific steps of this process are detailed as follows: Upper branch: For perform depthwise separable convolution with a convolution kernel size fixed at 3x3 Lower branch: The lower branch of the model uses random channels to generate sub-features and gradually captures multi-scale semantic responses incrementally. In this specific branch. The input features are divided into two groups according to the channel size: and This separation is determined by using two different formulas: A j,1 ∈R H×w×C and A j,2 ∈R H×w×CThe dilated convolution is followed by the FRN-ARelu activation function to suppress the abnormal propagation of gradients and finally by a 1x1 convolution. Therefore, combining these two components can form the attention map A j For each subgroup, channel dependencies are thus captured.
[0025] To address the issues of positive-negative sample imbalance in the damaged building detection task and the vanishing regression gradient of the CloU loss function when dealing with small targets, this paper proposes the T-Focal loss function. This function combines the geometric alignment constraint of CloU and the distribution matching property of NWD: CloU optimizes the bounding box regression accuracy of regular-scale targets through the center distance penalty term and the aspect ratio similarity metric, while NWD constructs a probability distribution matching mechanism for the positions of small targets based on the Wasserstein distance and reduces the coordinate sensitivity through Gaussian modeling. To address the gradient drowning effect caused by sample imbalance. NWD measures the similarity between the predicted frame and the ground truth frame using the Wasserstein distance and has better robustness for small target detection, especially suitable for handling targets of different scales. The Wasserstein distance first models the predicted frame and the ground truth frame as two-dimensional Gaussian distributions The mean of each distribution is the center point of the bounding box, and the variances are the width and height of the box, defined as follows: c a =(c xa , c ya ) and c b =(c xb , c yb ) represent the center coordinates of the predicted frame and the ground truth frame respectively. The Wasserstein distance provides a more accurate measure of the similarity between the predicted box and the ground truth box. However, as a distance metric, it cannot be directly applied to quantify the similarity of bounding boxes. The NWD loss is normalized and defined as follows: Here C is a dataset-specific constant. The NWD loss is finally defined as follows: The Tversky Loss dynamic adjustment mechanism is introduced to emphasize its complementary role in strengthening the compensation for class imbalance. The Tversky Loss function controls the balance between false positives FP and false negatives FN by introducing weights. It is particularly effective for imbalanced datasets. The formula is as follows: TP is the true positive, FP represents the false positive, FN stands for the false negative, and δ is a small constant used to avoid division by zero. In this context, λ acts as the weight for false negatives, while β specifies the penalty for false positives. This parameterization allows the model to explore deeper damaged regions during training. The final formula is α is a weight parameter used to adjust the share of CloU and NWD in the bounding box regression loss. The four of them construct an adaptive multi-objective optimization mechanism by dynamically adjusting the weight ratios of each loss term. This method combines the collaborative optimization of multi-structural feature extraction networks, cross-scale attention modules, and combined loss functions, effectively improving the feature discrimination of damaged targets in complex scenarios, while maintaining the integrity of target shape details and enhancing the performance of the model.
[0026] The experimental environment for this experiment is as follows: Processor Intel Core i9-12900H; Memory is 64GB DDR4 (3200MHZ); Graphics card GeForce RTX 3090; Storage hard drive is 1024GB PCIe 4.0 NVMe solid-state drive; Operating system is Ubuntu 20.04 Development environment: Python 3.9.17 pytorch 2.0.1 torchvision 0.15.2
[0027] In summary, the present invention proposes a remote sensing damaged building target detection algorithm based on YOLOv8. Aiming at the problems of target misdetection, missed detection, and excessive regression loss caused by factors such as complex backgrounds, small target sizes, dense distributions, and variable damage degrees in remote sensing images of damaged buildings, we designed a new improved model. This model provides an innovative detection method that can significantly improve the detection accuracy and efficiency of remote sensing damaged building images, has high practical application value, and is especially suitable for the task of detecting damaged building targets.
[0028] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of this invention patent should be subject to the appended claims.
Claims
1. A method for detecting damaged building targets in remote sensing based on YOLOv8, characterized in that, Including: Obtain and preprocess the remotely sensed damaged building dataset, clean and annotate the images, and then divide them into a training set, a validation set, and a test set. Among them, the remotely sensed damaged building image dataset includes remotely sensed damaged building images under different disaster conditions; Replace the original convolution module of YOLOv8 with the CDFR module to enhance the small target feature extraction ability; add the EDCA module to the neck end of YOLOv8 and the Concat of the Head to optimize the cross-scale feature fusion through dilated convolution and gated weighting mechanism; adopt T-Focalloss as the loss function, and fuse the geometric alignment, distribution matching, and sample balancing mechanisms; Use the training set to iteratively train the improved model to generate target weights; Use the test set to verify the model performance and output the detection results of damaged buildings.
2. The remote sensing damaged building target detection algorithm according to claim 1, characterized in that, After inputting the test set into the target network model for testing, it further includes: evaluating the performance of the tested target network model through preset performance evaluation indicators.
3. The remote sensing damaged building target detection algorithm according to claim 1, characterized in that, The CDFR module includes: 3×3 standard convolution and 1×1 point convolution deployed in parallel in the front-stage branch for multi-scale basic feature extraction. The rear-stage branch integrates a dual-path deformable convolution module to optimize the capture of geometric deformation features through dynamic sampling point offset. Filter Response Normalization (FRN) and Adaptive Rectified Linear Unit (ARelu) are used as activation functions to suppress the abnormal propagation of gradients, and the five-level feature maps are fused through channel dimension splicing and the output dimension is H / 2*W / 2*C.
4. The remote sensing damaged building target detection algorithm according to claim 1, wherein The EDCA module is deployed at the front end of the upsampling layer of the feature pyramid and the back end of the feature fusion layer of the detection head. It expands the receptive field through dilated convolution to capture global context information, integrates a cross-dimensional attention gating mechanism to dynamically reconstruct the channel weight distribution, retains the spatial details of small targets before upsampling and optimizes the cross-scale semantic combination after feature fusion. SoftMax is used to generate the channel attention heat map and combined with Channel Shuffle to promote cross-group feature migration.
5. The remote sensing damaged building target detection algorithm according to claim 1, wherein The T-Focalloss is composed of the CloU loss function, the Normalized Wasserstein Distance (NWD) loss function, and the dynamic balance mechanism of Tversky Loss and Focal Loss. Among them, the CloU loss optimizes the geometric alignment of the bounding box through the center point distance penalty term and the aspect ratio similarity metric. The NWD loss models the robustness of the small target position based on the Gaussian distribution. The Tversky Loss strengthens the penalty of false negative samples through adjustable parameters and the Focal Loss suppresses the contribution of easy-to-classify samples. The total loss formula is: Where α, λ, β are adjustment coefficients, and δ is a smoothing constant.
6. The remote sensing damaged building target detection algorithm according to claim 3, characterized in that The deformable convolution module optimizes the sampling positions of shallow details and deep semantic features respectively through a dynamic offset learning strategy to improve the feature representation ability of damaged buildings with irregular shapes.
7. The remote sensing damaged building target detection algorithm according to claim 3, characterized in that, The dual-path feature enhancement strategy of the EDCA module is that the upper branch uses depthwise separable convolution (3×3 kernel) to extract local features, and the lower branch generates multi-scale semantic responses through dilated convolution and uses FRN-ARelu activation to suppress gradient anomalies. Finally, the output features of the two branches are fused by element-wise addition (⊕).
8. The remote sensing damaged building target detection algorithm according to claim 4, characterized in that, The dual-path feature enhancement strategy of the EDCA module is that the upper branch uses depthwise separable convolution (3×3 kernel) to extract local features, and the lower branch generates multi-scale semantic responses through dual-channel dilated convolution and uses FRN-ARelu activation to suppress gradient anomalies. Finally, the output features of the two branches are fused by element-wise addition (⊕).
9. The remote sensing damaged building target detection algorithm according to claim 5, characterized in that, The loss function module T-Focal loss replaces the original CioU module in the detection network of YOLO v8 as the bounding box regression loss function.
10. The infrared image aircraft target detection method according to claim 1, wherein The inputting the training set into the network model for iterative training includes: Inputting the training set into the network model, and performing multiple rounds of iterative training on the network model through the SGD optimizer until the network model converges to obtain the target network model.
Citation Information
Cited By
Aerial bird family identification method based on deep learning
CN121617127A